Could one or two huge categories be skewing this whole result?
Checking whether the flat naming pattern across category sizes depends on a few outsized groups, or whether it holds because no single group dominates the sample.
largest_group_share_pct
Whenever a study pools data from groups of very different sizes, there is a standard worry: maybe the aggregate result is really just describing one or two dominant groups, with everything else along for the ride. If, say, a single enormous category made up half the sample, its naming behavior alone could set the overall 91% figure, and the apparent consistency across size segments could be an illusion created by that one group's weight rather than a genuine pattern shared across many categories.
This study can rule that specific worry out with a direct check. Across 193 groups and 752,124 rows, the single largest group accounted for only 1.5% of all rows in the sample. That is a small share for a maximum. It means no individual category, however large in absolute terms, is anywhere close to dominating the pooled totals reported elsewhere in this study, such as the 684,307 of 752,124 answer sets that named six or more businesses. The overall figures reflect a genuine aggregate across many groups of comparable weight, not one or two outsized contributors.
This is also why the size-segment analysis is trustworthy as a comparison of like with like. The four segments, from a smallest quarter of 88,655 rows to a largest quarter of 261,745 rows, are each built from many separate groups rather than one or two big ones standing in for a whole quarter. If the largest quarter's 261,745 rows were actually just one giant category, its 90% six-or-more rate would be a fact about that one category, not about "large categories" as a class. Because the maximum group share is only 1.5%, that quarter's total is necessarily built from a broad spread of large-but-comparable groups, and the 90% figure describes that spread rather than a single case.
The same logic applies to data quality more broadly. Before analysis, 3 corrections were applied to the dataset, and 0 rows were excluded out of 752,124 total, meaning the full dataset was usable without needing to throw out a meaningful fraction of it. A study that has to exclude a lot of rows to reach a clean result invites the question of whether the excluded rows would have told a different story. Here, that question does not arise in a serious way, since exclusions were zero and the corrections applied were few.
Put together, these facts support treating the study's central claim, that the six-or-more naming share stays close to constant from the smallest quarter (93.2%) to the largest (90%), as a genuine pattern across a broad and fairly even set of categories, rather than an average dominated by a small number of unusually large or unusually influential groups. The flatness observed is not a statistical accident produced by one big category; it is a pattern that shows up broadly across a sample where no single group holds more than 1.5% of the weight.
Other findings
When an AI engine answers a question about a small, niche market category, does it name fewer distinct businesses than it would for a huge, well-known category?
Do small categories get fewer named businesses than large ones?
A look at whether category size predicts how many distinct businesses an AI engine lists in its answer, across four size segments spanning from the smallest to the largest quarter of the sample.
Does the share of answers naming many businesses rise steadily as category size increases, or is it flat across all size groups?
How consistent is the naming pattern across all four size segments?
Checking whether the six-or-more-names share moves in a gradient from small to large categories, or stays essentially level across all four segments.
Could the overall flat result be an average that hides wide variation between individual categories, some naming almost everyone and others naming almost no one?
Does the flat pattern hold up within individual categories, not just across size groups?
Checking the spread of the six-or-more-names rate across the 188 individual category groups measured, to see whether the aggregate figure is representative or an average masking large swings.
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
- Your category measured the same way
- Your own numbers, not a sample deck
- Fifteen minutes, no obligation
