Do small categories get fewer named businesses than large ones?
A look at whether category size predicts how many distinct businesses an AI engine lists in its answer, across four size segments spanning from the smallest to the largest quarter of the sample.
by_segment.1 smallest quarter.classes.names 6 or more.pct
A reasonable guess, before looking at any data, is that AI engines behave like a thin encyclopedia for small categories and a thick one for big categories. A huge category, something like coffee shops or accounting software, has thousands of real businesses competing for attention, a long documented history, and presumably a deep well of named entities for a language model to draw on. A small category, a niche tool or a regional service type, has fewer real businesses to draw on in the first place. So it would be unsurprising if engines named fewer distinct businesses when answering about small categories, simply because there is less to name.
This study tested that guess directly. It took 193 groups covering 752,124 rows of answer data (none excluded), split every category into one of four size segments by row count, from a smallest quarter of 88,655 rows up to a largest quarter of 261,745 rows, and classed every answer set as naming six or more distinct businesses or naming five or fewer. If category size drove naming behavior, the smallest quarter should show a noticeably lower share of six-or-more answers than the largest quarter.
It does not. In the smallest quarter, 93.2% of answer sets named six or more distinct businesses, against 6.8% that named five or fewer, out of 88,655 rows in that segment. That is barely different from the pattern in the whole sample, where 684,307 of 752,124 answer sets (91%) named six or more businesses. Small categories are not thin on names. Whatever mechanism produces a rich, multi-name answer does not appear to depend on how many real competitors exist in the category.
Why would this be true? The likely explanation is that an AI engine is not enumerating a market from a private census of real businesses, it is generating a plausible-sounding list from whatever patterns are strongest in its training data and retrieval context for that kind of question. A prompt asking "what are some good options for X" pulls a similar list-shaped response whether X is a huge category or a small one, because the shape of the answer is driven by the shape of the question and the model's general habit of listing several options, not by an internal inventory of how many businesses actually exist. A model does not know, in any strong sense, how big the category is before it starts answering. It infers a plausible list size from the pattern of similar prompts it has seen, and that pattern is fairly stable across category sizes.
An alternative explanation worth naming: maybe the size segments in this study are not tracking real-world category size well, so the smallest quarter here is not actually made of niche markets. That is worth checking, but it would take a different kind of study, one that validates segment size against an independent measure of real business counts. What this study can say is narrower and still useful: within the row-count based segmentation used here, the smallest and largest quarters produced almost identical naming behavior, 93.2% versus 90%.
For a company in a small or niche category, the practical implication is that being small is not a reason to expect thin, generic answers from AI engines. The engine is not holding back names because your category is small. If your business is not showing up in a six-or-more-name answer, category size is not the excuse. Something else, likely how the business is described, linked, or reviewed online, is more likely responsible than the size of the market it competes in.
Other findings
Does the share of answers naming many businesses rise steadily as category size increases, or is it flat across all size groups?
How consistent is the naming pattern across all four size segments?
Checking whether the six-or-more-names share moves in a gradient from small to large categories, or stays essentially level across all four segments.
Could the overall flat result be an average that hides wide variation between individual categories, some naming almost everyone and others naming almost no one?
Does the flat pattern hold up within individual categories, not just across size groups?
Checking the spread of the six-or-more-names rate across the 188 individual category groups measured, to see whether the aggregate figure is representative or an average masking large swings.
With market sizes varying so much, could a handful of dominant categories be driving the aggregate numbers rather than the pattern holding broadly?
Could one or two huge categories be skewing this whole result?
Checking whether the flat naming pattern across category sizes depends on a few outsized groups, or whether it holds because no single group dominates the sample.
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
- Your category measured the same way
- Your own numbers, not a sample deck
- Fifteen minutes, no obligation
