Methodology
Small categories get just as many distinct businesses named as huge ones, once you control for run count
Engines and exact versions
A study that does not name the model version it ran against is not reproducible, because the answer changes when the model does.
| Engine | Model version |
|---|
Run window, UTC: [object Object].
Method
The unit of analysis is a single answer set: one engine's response to one prompt about one category, reduced to a count of distinct businesses named. Every one of the 752,124 answer sets in this study was classed into exactly one of two buckets: 'names 6 or more' or 'names 5 or fewer.' The cut at six was fixed before looking at results and applied uniformly, so a category that got exactly six names on one occasion and five on another crosses the line, but the boundary itself does not move to flatter any group.
Categories, not individual businesses, are the thing being sized. Each category was assigned a row count reflecting its underlying content or query volume, and then all 193 category-groups were sorted by that count and cut into four segments of roughly equal population: a smallest quarter of 88,655 rows, a second quarter of 179,165 rows, a third quarter of 222,559 rows, and a largest quarter of 261,745 rows. This quartering by volume, rather than by an arbitrary threshold, means each segment represents a real quarter of the observed population rather than a hand-picked slice.
Two design choices are worth naming. First, the study did not exclude any rows: 0 of 752,124 rows were dropped, so the segments are not shaped by a hidden filter. Second, 3 corrections were applied before analysis, addressing known data issues such as duplicate answer sets or mis-parsed business names, without touching the six-name threshold itself. The average category-group in this dataset spans 3,897.02 answer-set units, and no single group dominates the total: the largest group accounts for only 1.5% of all rows, which matters for the next section because it means the headline split is not just one enormous category's behavior wearing a population-wide disguise.
Limitations we volunteer
Written by us, before anyone else found them.
- Single pass. Run-to-run variance is not characterised.
- Gemini's cited sources are largely unavailable through Google's API, so source analysis rests on the other engines.
Terms used in this study
- Class
- One of the two outcome buckets a given answer set falls into: naming six or more distinct businesses, or naming five or fewer.
- Segment
- One of four quarters that categories are sorted into by size, from the smallest quarter of categories to the largest, used to test whether category size predicts naming breadth.
- Group
- An individual category or query cluster in the dataset; there are 193 of them, and no single group holds more than 1.5% of all rows.
- Row
- One recorded answer instance in the dataset; the study covers 752,124 rows in total with 0 excluded from analysis.
- Distinct businesses named
- The count of separate, non-repeated business names an engine's answer produced for a given query, which is the raw measure before it gets classed as six-or-more or five-or-fewer.
- Class spread across groups
- The range, from minimum to maximum percentage, of how often a given class appeared when measured separately within each of the 188 groups, showing how much variation is hidden inside the segment-level averages.
- Mitigation
- A correction applied to the raw data before analysis, such as removing duplicate or malformed entries; 3 such corrections were applied in this study.
- Measurement window
- The calendar period during which answer data was collected for this study, running from 2026-04-27 to 2026-08-16.
References
Sources this study reads against. Every link was fetched and confirmed reachable at publication.
- Industry Classification Overview U.S. Bureau of Labor Statistics, 2024 Describes how NAICS assigns establishments to detailed industry codes, relevant to noting we did not map our categories onto this taxonomy.
- Dropping diversity of products of large US firms: Models and measures arXiv, 2021 Discusses SIC-code industry grouping and within-industry similarity as a diversity-adjacent measure, cited for classification-method context.
- North American Industry Classification System (NAICS) U.S. Census Bureau, 2024 Defines the standard federal classification of business establishments, used as context for how categories are normally defined versus our own segmentation.
- Scaling and universality in urban economic diversification (open access) PMC (National Institutes of Health), 2016 Open-access copy of Youn et al., used for the specific scaling-exponent figures cited.
- Zipf distribution of U.S. firm sizes PubMed, 2001 Bibliographic record for Axtell 2001, cited alongside the full-text PDF.
- Scaling and universality in urban economic diversification Journal of the Royal Society Interface, 2016 Shows business-category diversity in cities scales sublinearly with population, checked across many metro areas, the closest published analogue to our category-size-vs-diversity finding.
- Demystifying entrepreneurial name choice: insights from the US biotech industry New England Journal of Entrepreneurship (Emerald), 2022 Uses the Blau index to measure name/category diversity, an alternative continuous method we note but do not apply.
- Zipf law and the firm size distribution: a critical discussion of popular estimators IDEAS/RePEc (Journal of Evolutionary Economics), 2015 Provides a critical re-examination of how robust the Zipf firm-size finding is to estimation method, used to caution against over-reading the Axtell benchmark.
- Zipf Distribution of U.S. Firm Sizes Science (author-posted PDF), 2001 Establishes that US firm sizes follow a Zipf distribution using the full population of tax-paying firms, the size-skew benchmark our category-diversity finding is contrasted against.
Data
The complete row-level dataset is published open and ungated under CC BY 4.0. Every number in this study can be recomputed from it.
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
- Your category measured the same way
- Your own numbers, not a sample deck
- Fifteen minutes, no obligation
