Methodology
After correcting for Anthropic's duplicate rows, how many distinct businesses does Claude actually name per buyer question, and does it differ from OpenAI or Gemini
Engines and exact versions
A study that does not name the model version it ran against is not reproducible, because the answer changes when the model does.
| Engine | Model version |
|---|
Run window, UTC: [object Object].
Method
The unit of analysis here is a single answer: one response from one engine to one query, evaluated once. Across the study, 926,792 such units were collected, distributed across 193 query groups, with an average of 4,802.03 units per group. A query group is a cluster of related prompts, for instance different phrasings of a request for a recommendation in the same category, run across the 3 engines being compared (ChatGPT, Claude, and Gemini) so that each group produces a roughly comparable set of answers per engine.
For each unit, the analysis counted the number of distinct, real business names that appeared in the answer text, then sorted that count into one of 4 classes:
| Class | What it means |
|---|---|
| names none | No business named at all |
| names 1 to 5 | A short, curated-feeling list |
| names 6 to 9 | A fuller list |
| names 10 or more | An extensive, near-exhaustive list |
This classing choice, rather than reporting a raw average count of names per answer, is deliberate. An average can be dragged around by a small number of extreme outliers, for instance a handful of answers that name thirty or forty businesses in a dump-style list. Classing puts every answer into a bucket regardless of exactly how extreme its tail is, and lets you see where the bulk of the mass actually sits. It also matches how a human reader experiences the answer: nobody reading a list of eleven names is thinking "eleven," they're thinking "a long list," and the class boundaries are set to track that qualitative jump.
A unit was counted as naming a business only when the answer used something identifiable as a proper business name, not a generic category description like "a local plumber" or "several options in your area." Answers that described options without naming any specific business landed in names none. This distinction is where a large share of the disagreement between casual impressions and careful counts tends to originate: an answer that feels information-dense because it describes many attributes can still name zero actual businesses, and an answer that feels short can still pack ten names into two sentences.
Limitations we volunteer
Written by us, before anyone else found them.
- Single pass. Run-to-run variance is not characterised.
- Gemini's cited sources are largely unavailable through Google's API, so source analysis rests on the other engines.
Terms used in this study
- Unit
- One answer given by one engine to one prompt. Each unit is classified into exactly one names-count class based on how many distinct businesses it mentions.
- Class
- A bucket describing how many businesses a single answer names: names none, names 1 to 5, names 6 to 9, or names 10 or more. Every unit falls into exactly one class.
- Segment
- In this study, the engine that produced the answer: ChatGPT, Claude, or Gemini. Class shares are reported both pooled across all segments and broken out per segment.
- Group
- A cluster of related query units, such as answers gathered under one prompt topic or category, used to check whether class shares hold steady across different subject matter rather than being an artifact of one topic.
- Largest group share
- The proportion of all units contributed by the single biggest group, used to confirm that no one query topic is large enough to distort the overall or per-engine class percentages.
- Mitigation
- A data-cleaning or deduplication step applied to the raw rows before classification, intended to prevent repeated or malformed rows from inflating a business's name count within a single answer.
- Excluded rows
- Rows removed from the dataset before analysis, for reasons such as malformed data. In this study that count is zero, meaning every collected row was classified and counted.
References
Sources this study reads against. Every link was fetched and confirmed reachable at publication.
- How Large Language Models Source Brand Reputation Across Languages and Markets arXiv (2606.25787) Survey-style paper linking to Peec AI and Omniscient Digital citation-sourcing studies, used here as a bibliography pointer rather than a primary figure source.
- The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages arXiv (2606.23165) Evidence that AI-mediated brand reputation varies systematically by language and market, supporting the case for checking results within groups rather than pooling.
- The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries arXiv (2601.00912) Adjacent benchmark on structural gaps in LLM brand/entity naming across architectures, cited for its finding that GEO optimization did not correlate with discovery success.
- Cultural Encoding in Large Language Models: The Existence Gap in AI-Mediated Brand Discovery arXiv (2601.00869) Group-level benchmark finding a 30.6 percentage point gap in brand mention rates between Chinese and International LLMs on identical English queries, the closest published precedent for measuring mention behavior by model group rather than pooling.
- An Analysis of AI Overview Brand Visibility Factors (75K Brands Studied) Ahrefs Primary industry dataset analyzing millions of AI Overview responses via Ahrefs Brand Radar, finding roughly 26% of the 75,000 brands studied had zero mentions, used as a comparison point for our own names-none class.
- Which Economic Tasks are Performed with AI? Anthropic Underlying paper describing the Clio-based task classification methodology in full, cited as the primary source for the O*NET-based grouping approach.
- Across 75,000 Brands, YouTube Mentions Are the Strongest Signal of AI Visibility, New Ahrefs Report Reveals Businesswire Press release giving the underlying dataset scale for the Ahrefs Q1 2026 report, including 146 million search result pages and 730,000 AI responses analyzed.
- Anthropic Economic Index report: Cadences Anthropic Extends the classification method to per-conversation artifact type on Claude's chat and Cowork surfaces, aggregated monthly, showing how Anthropic itself groups output rather than pooling it.
- Introducing the Anthropic Economic Index Anthropic Anthropic's own primary documentation of how Claude.ai conversations are classified using Clio and mapped to O*NET occupational tasks, the reference standard for how classification of Claude output is normally done.
- How to Track AI Search Engine Citations & Sources: The Complete Guide for 2026 Otterly.ai Industry-standard operational definitions distinguishing a brand mention (named, no link) from a citation (attributed with a link), used to scope what our study counted.
- The AI Citation Economy: What 1+ Million Data Points Reveal About Visibility in 2026 Otterly.ai Large-scale industry benchmark of over one million website citations across AI search platforms, reporting citation share by engine, used as a scale comparison for our per-answer counting approach.
- ChatGPT Gemini Claude Brand Mentions: Tracking Guide MaxAEO Secondary summary of the arXiv Existence Gap paper's group-level mention-rate effect, cited for its restatement of the 1,909-query, six-LLM, 30-brand design.
- Claude AI Brand Recommendations: How They Differ From ChatGPT and Perplexity MaxAEO Independent 1,400-prompt study of 50 B2B SaaS companies reporting that ChatGPT and Gemini mentioned 100% of tested brands versus 88% for Claude, and that Claude's shortlist runs roughly half the length of ChatGPT's, a coverage-rate finding this study's per-answer count distribution is checked against.
Data
The complete row-level dataset is published open and ungated under CC BY 4.0. Every number in this study can be recomputed from it.
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
- Your category measured the same way
- Your own numbers, not a sample deck
- Fifteen minutes, no obligation
