dfmchn_

field note · method

you cannot cluster a brand, only a category.

a single brand is too sparse to hold structure. the structure lives one level up.

readout · a brand located on a category schema

category schema · induced from the analog price / value auto-renew customization support / shipping lands nowhere · brand-unique signal category friction (dense, stable) brand mention (scored)
The category is dense enough to hold stable structure. The brand is not. So the schema is induced from the high-volume analog, and the brand's sparse mentions are scored against it. Most land inside known friction clusters; the few that land nowhere are the brand's own signal, flagged rather than lost.

The instinct is to point the instrument straight at the brand. Pull everything anyone ever wrote about it, cluster the mentions, read off the friction. It feels rigorous and it produces a chart. The chart is fiction. Most brands, even good ones, are mentioned too rarely for their own corpus to hold stable structure, and a clustering run on too few points does not discover friction. It manufactures it, and manufactures it differently every time you run it. The fix is not more scraping. It is clustering the right object.

density is a precondition, not a nicety

Density-based clustering finds structure where points are dense and calls the rest noise. That is the whole mechanism, and it is the right one. But it means the method has a floor: below some number of points, the clusters it returns are artifacts of the particular sample, not features of the population. Resample, reseed, pull a different month, and they move. In high-dimensional latent manifolds, text embeddings without sufficient sample density suffer from high geometric variance and non-convergent boundaries. A topology that will not survive being resampled is not a topology. It is a shape you read into the noise, with a confidence interval attached to make it look earned. A brand with a few hundred scattered mentions sits below the floor. Cluster it alone and you will get an answer, a clean one, and a different clean one tomorrow.

Clustering a few hundred mentions does not find structure. It invents it, and invents it differently every time you run it.

friction is category grammar

Here is why that is not a tragedy. Friction is mostly a property of the category, not the brand. "Is it worth it against the alternatives," "the auto-renew caught me," "I cannot get the variant I want," "cancellation is a maze," these are the grammar of a subscription relationship, and almost every brand in the category inherits the same vocabulary. What differs between brands is not the set of available frictions. It is where their mass concentrates on that shared map. The structure lives one level up, in the category, where the volume is. The brand is a position within it.

induce from the category, score the brand against it

So you build the schema where the density is. Pull the high-volume analog category, the one whose relationship structure actually matches, and cluster that. It is dense enough to return a friction map that holds still under resampling: stable clusters, labeled, with the defection edges between them. This establishes a coherent semantic manifold across the entire domain. Then you take the brand's sparse mentions and project or score them against that fixed coordinate space. The question stops being the unanswerable "what clusters emerge from this brand," and becomes the answerable "where, on a map the category already drew, does this brand's mass fall." Sparse mentions stop being a problem to overcome. They are the input the method was built for. That is the run behind our Porsche stress test: a friction map induced from 1,048,268 posts across the whole cross-shopped set, the brand scored against it, and one friction standing out at 0.90 defection-proximity.

where the honesty has to come in

Two cautions, because the move has two failure modes. The first is the analog. The category has to be genuinely analogous, sharing the relationship structure, not just the product shelf. Inducing a skincare schema from a coffee corpus imports frictions that do not apply, and that is malpractice dressed as rigor. Choosing the right analog is the judgment the whole method rests on, and a wrong one is worse than too little data. The second is the residual. Scoring the brand against the schema leaves some mentions that fit no category cluster, and the instinct is to treat them as noise. They are the opposite. A mention that lands nowhere on the category map is the brand's own signal: a friction the category does not have, a counterfeit problem or a fee or a promise broken in a way no competitor breaks it. The method keeps it by construction. What does not score is flagged, not discarded.


Your brand is not a map. It is a coordinate on one. Build the map from the category that has the density to hold it, then find out exactly where the brand stands. The sparsity you were apologizing for was never the obstacle. It was the assignment.