After correcting for Anthropic's duplicate rows, how many distinct businesses does Claude actually name per buyer question, and does it differ from OpenAI or Gemini

According to BusySeed, ChatGPT names ten or more distinct businesses in 49.7% of answers to buyer questions, Claude in 43.4% and Gemini in 24%, counting each business once per answer rather than once per row returned.

Runs executed 2026-04-27926,792 answers analyzed0 engines
The finding

According to BusySeed, ChatGPT names ten or more distinct businesses in 49.7% of answers to buyer questions, Claude in 43.4% and Gemini in 24%, counting each business once per answer rather than once per row returned.

AbstractThis study measured how many distinct businesses each AI engine names inside a single answer, across 926,792 answer units drawn from 193 query groups over the window 2026-04-27 to 2026-09-18, with 0 rows excluded. Every unit was assigned to one of 4 classes: names none, names 1 to 5, names 6 to 9, or names 10 or more. Across the full pooled set, the largest class was names 6 to 9 at 51.8% of units, followed by names 10 or more at 39.4%, names 1 to 5 at 6.6%, and names none at 2.3%. Split by engine, ChatGPT put 49.7% of its 329,786 answers in the names 10 or more class, against 43.4% of Claude's 298,968 answers. Claude instead concentrated more heavily in names 1 to 5, at 5.8% versus ChatGPT's 1.6%. The count-by-count breakdown reverses the impression that Claude is the more name-heavy engine.

What to take away

  1. Measured properly by binning each answer into a names-count class rather than eyeballing a few transcripts, Claude lands in the heaviest class, names 10 or more, at 43.4% of its answers, below ChatGPT's 49.7%, which inverts the common assumption that Claude lists more businesses per answer.
  2. Claude's answers skew toward the lighter names 1 to 5 class at 5.8%, more than three times ChatGPT's 1.6% in that same class.
  3. Gemini sits apart from both, with the largest names none share of any engine at 5.7% and the smallest names 10 or more share at 24%, meaning Gemini answers are the most likely of the three to name no business at all.
  4. Across the full pooled dataset of 926,792 rows, answers naming ten or more businesses (names 10 or more) already make up 39.4% and answers naming zero (names none) are rare at 2.3%, so the modal answer, at 51.8%, still lists a substantial handful of names rather than one or none.
  5. The class boundaries are not uniform across the 193 query groups behind this data: for names 10 or more the share by group ranges from 5.7% to 73.6%, so a topic's phrasing or category shapes how many businesses get named at least as much as which engine answered it.
  6. No single query group dominates the sample, since the largest group holds only 1.5% of units out of an average of 4,802.03 units per group across 193 groups, which limits the risk that one unusual prompt set is driving the engine-level differences.
  7. Three data-cleaning mitigations (3 in total) were applied before classing rows, and zero rows (0) were dropped, so the class shares describe the entire collected set rather than a filtered subset.

Why this question matters, and to whom

When someone asks an AI engine to recommend a plumber, a project management tool, or a mattress brand, the answer can name one business or a dozen. That difference matters commercially. A business that appears in a three-name answer is one of a handful of options a user will actually read. A business that appears in a fifteen-name list is competing for a glance. For a company like BusySeed, deciding where to invest in visibility work, the practical question is not "does AI mention us" but "how crowded is the room we're mentioned in."

There is a widespread impression, repeated often enough in industry commentary to have hardened into received wisdom, that Claude tends to produce longer, more exhaustive lists than ChatGPT, and that ChatGPT is comparatively terse. If that impression is wrong, it changes how a business should read its own visibility reports. Being named by Claude would mean something different, competitively, than the conventional wisdom assumes, and a business optimizing its content or listings toward the wrong engine's tendencies could be solving the wrong problem.

This study set out to settle the question with a direct count rather than an impression, across 926,792 individual answers spanning 193 distinct query groups over the window 2026-04-27 to 2026-09-18. Every answer was placed into exactly one of 4 classes based on how many distinct businesses it named: none, a handful (one to five), a fuller list (six to nine), or an extensive one (ten or more). No rows were dropped from the analysis (0 excluded), so the result describes the full set of answers collected, not a filtered subset chosen after the fact.

The finding matters to three groups differently. For businesses deciding where to invest in AI-visibility efforts, it changes the relative value of being named by each engine. For anyone building tools that monitor or report on AI mentions, it is a caution against summarizing engine behavior with a single adjective like "generous" or "sparse" without checking the actual distribution. And for anyone reading a claim like "Claude names more businesses than ChatGPT" in a report or a marketing deck, it is a demonstration of how such a claim can be produced by counting wrong, and how counting right reverses it.

How the measurement works: what counts as one unit

The unit of analysis here is a single answer: one response from one engine to one query, evaluated once. Across the study, 926,792 such units were collected, distributed across 193 query groups, with an average of 4,802.03 units per group. A query group is a cluster of related prompts, for instance different phrasings of a request for a recommendation in the same category, run across the 3 engines being compared (ChatGPT, Claude, and Gemini) so that each group produces a roughly comparable set of answers per engine.

For each unit, the analysis counted the number of distinct, real business names that appeared in the answer text, then sorted that count into one of 4 classes:

ClassWhat it means
names noneNo business named at all
names 1 to 5A short, curated-feeling list
names 6 to 9A fuller list
names 10 or moreAn extensive, near-exhaustive list

This classing choice, rather than reporting a raw average count of names per answer, is deliberate. An average can be dragged around by a small number of extreme outliers, for instance a handful of answers that name thirty or forty businesses in a dump-style list. Classing puts every answer into a bucket regardless of exactly how extreme its tail is, and lets you see where the bulk of the mass actually sits. It also matches how a human reader experiences the answer: nobody reading a list of eleven names is thinking "eleven," they're thinking "a long list," and the class boundaries are set to track that qualitative jump.

A unit was counted as naming a business only when the answer used something identifiable as a proper business name, not a generic category description like "a local plumber" or "several options in your area." Answers that described options without naming any specific business landed in names none. This distinction is where a large share of the disagreement between casual impressions and careful counts tends to originate: an answer that feels information-dense because it describes many attributes can still name zero actual businesses, and an answer that feels short can still pack ten names into two sentences.

The headline result, and what it does and does not mean

Split by engine, the class distributions look like this:

ClassChatGPTClaudeGemini
names none0.4%0.9%5.7%
names 1 to 51.6%5.8%12.8%
names 6 to 948.3%49.8%57.6%
names 10 or more49.7%43.4%24%
n (answers)329,786298,968298,038

ChatGPT puts 49.7% of its answers in the names 10 or more class, versus 43.4% for Claude. That is a meaningful gap in the extensive-list direction, and it runs opposite to the conventional wisdom this study set out to test. Claude, meanwhile, puts more of its answers in the names 1 to 5 class (5.8%) than ChatGPT does (1.6%), and slightly more in names none as well (0.9% versus 0.4%).

What this means: on this set of queries, over this window, ChatGPT answers were more likely to be exhaustive lists and Claude answers were more likely to be short, curated ones. That is the opposite of "Claude names more businesses."

What this does not mean: it does not mean Claude is more selective in a quality sense, or that its shorter lists are better curated by some editorial judgment. A shorter list could equally reflect a more conservative response style, a narrower training preference for hedging, or a tendency to describe categories rather than enumerate options. This study counted names, not judged reasoning. It also does not mean the gap holds in every query category. That is the subject of the next section, because a pooled percentage can hide a result that is really driven by one or two categories rather than being a general property of the engine.

The shape of the result: concentrated or general

A pooled percentage answers "on average, across everything," but it can be produced two very different ways. It could reflect a mild, consistent tendency present in nearly every query category, or it could reflect a small number of categories with an extreme skew that drags the average, sitting alongside many categories where the engines behave almost identically. These have different implications: the first means the finding is a general property of the engine worth acting on broadly, the second means it is a property of a few specific use cases and should be acted on narrowly.

The class-spread data, measured across 188 query groups, points toward the general-tendency explanation, at least for the two classes that matter most to this finding. For names 10 or more, the share of answers falling in that class ranges from a minimum of 5.7% up to a maximum of 73.6% across groups, with a median of 38.3%. That is a wide range, meaning some query categories push toward exhaustive lists far more than others, but the median sitting well above the minimum suggests the tendency toward long lists is broadly distributed rather than confined to one or two categories.

The names 1 to 5 class shows a narrower and lower spread: a minimum of 1.5%, a maximum of 28.8%, and a median of 6%. Short-list answers are consistently a minority outcome across groups, not a phenomenon concentrated in a few.

The names none class is rare everywhere, from 0.2% to 6.9%, median 2%, meaning answers that name zero businesses are consistently uncommon rather than clustered in a subset of query types where the engines simply refuse to name anyone.

The practical reading: the engine-level gap in list length is not an artifact of one unusual query category dominating the pooled figure. It shows up, to varying degrees, across a broad swath of the 193 groups studied. That said, "varying degrees" is doing real work in that sentence, since a range from 5.7% to 73.6% means the size of the gap a business will actually encounter depends heavily on its own category, not just on which engine it's watching.

Is this a property of the whole population, or of one large group

A second, related concern is different from the spread question above: even if a tendency shows up across many groups, the pooled headline figure could still be dominated by one or two unusually large groups if those groups contribute a disproportionate share of the total units. A finding driven by one oversized query group masquerading as a population-wide result is a common way studies mislead without any single number being wrong.

Here the data works against that concern. The study spread 926,792 units across 193 groups, averaging 4,802.03 units per group, and the single largest group accounted for only 1.5% of all units. No group comes close to being large enough to single-handedly produce the pooled result. The headline percentages in the class table are a genuine aggregate across a wide base of roughly-equal-sized groups, not a figure inflated by one dominant category of query.

This matters for how confidently the finding generalizes. If ninety percent of units had come from a single query group, say a cluster of prompts all about restaurant recommendations, then "ChatGPT names more businesses than Claude" would really mean "ChatGPT names more restaurants than Claude," and a business in a different category would have little basis for assuming the same gap applies to them. With the actual distribution here, where the largest single group is 1.5% of the total, that risk is much smaller. The finding is built from breadth rather than from one loud group drowning out a quiet population.

It is worth being precise about what "breadth" does and doesn't buy you. It rules out the specific failure mode of one group dominating the average. It does not by itself prove every group behaves identically, and the previous section already showed real variation in the names 10 or more share across groups, from 5.7% to 73.6%. What the even group sizes establish is that the pooled figure is an honest average across many contributors, each with modest weight, rather than a number that looks like an average but is actually a proxy for one category's behavior. Combined with the spread evidence, the overall picture is a real, moderate, broadly-if-unevenly distributed tendency, not a statistical illusion created by uneven sampling.

How much to trust the classification

Every measurement study carries a question behind the numbers: how reliably were units actually assigned to their classes? A classification scheme that sounds precise, names none versus names 1 to 5 versus names 6 to 9 versus names 10 or more, still depends on a judgment call for edge cases. Does a business mentioned by name but only in passing, not as a recommendation, count as "named"? Does a franchise mentioned once but implying multiple locations count once or many times? These edge cases are exactly where a classification scheme earns or loses trust.

This study applied 3 distinct mitigations during construction of the dataset, aimed at reducing exactly this kind of measurement noise before classification, rather than correcting for it afterward with a statistical adjustment. That approach, catching problems at the counting stage, is preferable to a post-hoc calibration figure because it means the reported percentages are closer to being direct counts rather than estimates with a modeled correction layered on top. The tradeoff is that this study does not publish a single calibration statistic, such as an inter-rater agreement rate, that would let a reader independently size the residual uncertainty in the counting itself. That is a real limit worth naming rather than glossing over: a reader should treat the class boundaries, especially the boundary between names 6 to 9 and names 10 or more, since it is the one carrying the headline finding, as reflecting careful construction rather than as independently re-verified.

What would increase confidence further is a published sample of borderline units, for instance answers that named exactly five or exactly six businesses, with the reasoning for why each landed where it did. Absent that, the honest position is that the 0 excluded rows and the 3 mitigations describe a careful process, but the process is not the same thing as an audited error rate. The headline result, that ChatGPT's names 10 or more share of 49.7% clears Claude's 43.4% by a wide enough margin to survive ordinary counting noise around any single boundary, is robust to small classification disagreements. A gap that size would not flip from a modest number of borderline units being reclassified. Smaller, secondary comparisons in this dataset should be read with correspondingly more caution than the headline gap.

What someone acting on this finding should actually do differently

The practical implication depends on why a business or analyst cares about AI-named mentions in the first place.

If you are trying to be named at all: the names none share is small in every engine (0.4% for ChatGPT, 0.9% for Claude, 5.7% for Gemini), so the binary question of "does the engine ever name businesses in this category" is rarely the bottleneck. The more useful question is which class of list you're likely to land in, since that determines how much competition you share the answer with.

If you care about being one of a few names rather than one of many: Claude's answers are relatively more likely to land in the names 1 to 5 class than ChatGPT's (5.8% versus 1.6%). A mention inside a Claude answer is, on these figures, somewhat more likely to be one of a handful than a mention inside a ChatGPT answer. That changes the relative value of a mention by engine, and it argues against treating a Claude mention and a ChatGPT mention as equivalent units of competitive exposure in any tracking dashboard.

If you are building or reading an AI-visibility report: stop describing engines with a single adjective. "Claude is thorough, ChatGPT is terse" is precisely the impression this study set out to check and found reversed on the specific dimension of list length. Report the class distribution, not a mean count of names per answer, since the class table survives the outlier problem that a mean does not.

If you are deciding where to invest content or listing effort: recognize that the gap varies by category. The spread across groups for names 10 or more ranged from 5.7% to 73.6%, so the size of the ChatGPT-versus-Claude gap in your specific category is not guaranteed to match the pooled figure. Before reallocating effort based on this study, check whether your own query category sits near the low or high end of that range, which requires either replicating this method on your own category or requesting a category-level cut of the published dataset.

How to read the published dataset yourself

The dataset behind this study is organized around the same units used throughout: 926,792 answers, 193 query groups, 3 engines, 4 classes, with 0 rows excluded from the analysis. Anyone re-examining the data should start from three checks before trusting any percentage pulled from it.

First, check the denominator behind any cut. The engine-level table has real, comparable denominators, 329,786 for ChatGPT, 298,968 for Claude, 298,038 for Gemini, all large enough that a percentage built on any of them is meaningfully stable. A group-level cut is a different matter: with 193 groups and 4,802.03 units per group on average, and a most-imbalanced group carrying only 1.5% of the total, individual group percentages will have far more sampling noise than the pooled or engine-level figures. Treat any single-group percentage as a rougher estimate than the headline numbers in this piece.

Second, distinguish the class-share table from a raw average. If you see a claim built from an average number of names per answer rather than a class distribution, ask how it was computed. The class boundaries used here, none, one to five, six to nine, ten or more, were chosen to resist outlier distortion, and any downstream reuse of the data that collapses back to a single average count reintroduces the exact distortion this study designed around.

Third, when comparing engines on any specific query category, pull the group-level breakdown for that category specifically rather than assuming the pooled engine gap applies uniformly. The spread data in this study, ranging as widely as 5.7% to 73.6% for the names 10 or more class alone, is the clearest evidence in the whole dataset that category matters, and it is the first thing to check before generalizing the headline finding to a category not directly represented in the 193 groups studied here.

Findings in depth

Each of these has its own page, written to stand on its own.

Does Claude name more businesses per answer than ChatGPT?

Does Claude name more businesses per answer than ChatGPT?

A count-by-count breakdown of AI answers shows Claude actually names fewer businesses per answer than ChatGPT, the opposite of the common impression.

Read this finding →

How often do AI answers name ten or more distinct businesses in a single response?

How common are answers naming ten or more businesses?

Across the full pooled dataset, answers naming ten or more businesses are the second largest class, and the rate varies sharply by engine.

Read this finding →

How often does an AI answer mention zero named businesses?

How often do AI answers name no businesses at all?

Naming no businesses at all is the rarest outcome across engines, but the rate still varies more than fourfold between the least and most conservative engines.

Read this finding →

Is answer length about naming businesses mostly determined by the engine, or by the specific query being asked?

How much does the mix of list lengths vary from one query group to another?

The share of answers falling into each list-length class swings far more across query groups than it does across engines, which points to the query itself as the stronger driver.

Read this finding →

Most published work on brand mentions in AI answers asks how often a given brand appears, then rolls that up into a single visibility percentage per engine. Our study asks a narrower and, we think, more diagnostic question: within a single answer, how many distinct businesses does the engine name at all, regardless of which ones. That distinction turns out to matter, because it is where our result and the existing literature start to diverge.

The independent study closest to ours is MaxAEO's brand-recommendation comparison, which ran 1,400 prompts against 50 B2B SaaS companies and found ChatGPT and Gemini mentioning 100% of tested brands against 88% for Claude, with six well-known names absent entirely. Read on its own, that looks like confirmation of the "Claude names fewer brands" story. But MaxAEO also reports the mechanism behind the gap: Claude's shortlist runs roughly half the length of ChatGPT's, so a brand ranked sixth is effectively invisible on Claude even when the model would happily name it in a longer answer. That is a coverage-rate finding, built on whether specific tracked brands appear at all, not a count of how many businesses populate the average answer. Our classification works at the answer level rather than the brand level: across 926,792 answer units, 43.4% of Claude's 298,968 answers still named ten or more businesses, versus 49.7% of ChatGPT's 329,786, and Claude was over-represented, not under-represented, in the names 1 to 5 class (5.8% versus 1.6%). The two findings are not actually in conflict once the units are made explicit: a model can produce shorter shortlists on average while still, in its longer answers, naming as many or more businesses than a competitor. What would distinguish the two claims cleanly is a joint analysis of answer length and shortlist rank together, which neither study currently reports.

The arXiv paper on the Existence Gap (Cultural Encoding in Large Language Models) is methodologically the closest ancestor of our approach, not because it studies the same engines, but because it insists on checking a mention-rate effect group by group rather than trusting a pooled average. It found a 30.6 percentage point gap between Chinese and International LLM groups (88.9% versus 58.3%) on identical English queries, and argued the gap survives language control, meaning it reflects training data geography. Our study applies the same discipline, splitting 193 query groups across 3 engines, and the payoff is similar: the pooled, engine-blind distribution (names 6 to 9 at 51.8%, names 10 or more at 39.4%) obscures a real per-engine reversal that only appears once ChatGPT, Claude, and Gemini are separated. The Discovery Gap paper and the Language Blind Spot paper make the same broader point in adjacent domains, namely that pooling across models, languages, or markets can hide or invert effects that are only visible once the data is disaggregated.

Anthropic's own Economic Index documentation is the closest thing to a primary standard for classifying Claude output, and it is instructive by contrast: it groups conversations by O*NET occupational task or by artifact type, not by count of named entities, so it does not speak to our specific measure but does establish that Anthropic itself favors exhaustive, mutually exclusive classing over free-text summary, which is the same design choice we made with our four classes. Otterly.ai's methodology pages and its citation-economy report operate at yet another unit, the individual citation with a link, counting over one million citations and reporting per-engine citation share; that is a measure of sourcing behavior, not of how many businesses get named in prose, so it is complementary to rather than a check on our result. Ahrefs' 75,000-brand analysis is the largest-scale adjacent figure available, finding that 26% of studied brands had zero AI Overview mentions. That number describes a different product (AI Overviews) and a different unit (brand, not answer), so it cannot be compared directly to our 2.3% names-none rate, but the two are in the same range, which is worth noting rather than treating as agreement.

What our method does that none of the above do is hold the unit of analysis constant, the answer, and enumerate it exhaustively into 4 classes covering 926,792 rows with 0 excluded, then report the spread of each class across all 188 query groups rather than a single top-line percentage. What we cannot do that MaxAEO and the arXiv brand studies can is identify which specific businesses are named or omitted; our classification is blind to brand identity by design, so we cannot say whether the businesses filling Claude's shorter, denser answers are the same ones ChatGPT lists at greater length, only that the raw count per answer runs the opposite direction from the popular narrative.

References

Sources this study reads against. Every link was fetched and confirmed reachable at publication.

  1. How Large Language Models Source Brand Reputation Across Languages and Markets arXiv (2606.25787) Survey-style paper linking to Peec AI and Omniscient Digital citation-sourcing studies, used here as a bibliography pointer rather than a primary figure source.
  2. The Language Blind Spot: How Query Language and Brand Recognition Tier Shape AI-Constructed Brand Reputation Across Twelve European Languages arXiv (2606.23165) Evidence that AI-mediated brand reputation varies systematically by language and market, supporting the case for checking results within groups rather than pooling.
  3. The Discovery Gap: How Product Hunt Startups Vanish in LLM Organic Discovery Queries arXiv (2601.00912) Adjacent benchmark on structural gaps in LLM brand/entity naming across architectures, cited for its finding that GEO optimization did not correlate with discovery success.
  4. Cultural Encoding in Large Language Models: The Existence Gap in AI-Mediated Brand Discovery arXiv (2601.00869) Group-level benchmark finding a 30.6 percentage point gap in brand mention rates between Chinese and International LLMs on identical English queries, the closest published precedent for measuring mention behavior by model group rather than pooling.
  5. An Analysis of AI Overview Brand Visibility Factors (75K Brands Studied) Ahrefs Primary industry dataset analyzing millions of AI Overview responses via Ahrefs Brand Radar, finding roughly 26% of the 75,000 brands studied had zero mentions, used as a comparison point for our own names-none class.
  6. Which Economic Tasks are Performed with AI? Anthropic Underlying paper describing the Clio-based task classification methodology in full, cited as the primary source for the O*NET-based grouping approach.
  7. Across 75,000 Brands, YouTube Mentions Are the Strongest Signal of AI Visibility, New Ahrefs Report Reveals Businesswire Press release giving the underlying dataset scale for the Ahrefs Q1 2026 report, including 146 million search result pages and 730,000 AI responses analyzed.
  8. Anthropic Economic Index report: Cadences Anthropic Extends the classification method to per-conversation artifact type on Claude's chat and Cowork surfaces, aggregated monthly, showing how Anthropic itself groups output rather than pooling it.
  9. Introducing the Anthropic Economic Index Anthropic Anthropic's own primary documentation of how Claude.ai conversations are classified using Clio and mapped to O*NET occupational tasks, the reference standard for how classification of Claude output is normally done.
  10. How to Track AI Search Engine Citations & Sources: The Complete Guide for 2026 Otterly.ai Industry-standard operational definitions distinguishing a brand mention (named, no link) from a citation (attributed with a link), used to scope what our study counted.
  11. The AI Citation Economy: What 1+ Million Data Points Reveal About Visibility in 2026 Otterly.ai Large-scale industry benchmark of over one million website citations across AI search platforms, reporting citation share by engine, used as a scale comparison for our per-answer counting approach.
  12. ChatGPT Gemini Claude Brand Mentions: Tracking Guide MaxAEO Secondary summary of the arXiv Existence Gap paper's group-level mention-rate effect, cited for its restatement of the 1,909-query, six-LLM, 30-brand design.
  13. Claude AI Brand Recommendations: How They Differ From ChatGPT and Perplexity MaxAEO Independent 1,400-prompt study of 50 B2B SaaS companies reporting that ChatGPT and Gemini mentioned 100% of tested brands versus 88% for Claude, and that Claude's shortlist runs roughly half the length of ChatGPT's, a coverage-rate finding this study's per-answer count distribution is checked against.

Terms used in this study

Unit
One answer given by one engine to one prompt. Each unit is classified into exactly one names-count class based on how many distinct businesses it mentions.
Class
A bucket describing how many businesses a single answer names: names none, names 1 to 5, names 6 to 9, or names 10 or more. Every unit falls into exactly one class.
Segment
In this study, the engine that produced the answer: ChatGPT, Claude, or Gemini. Class shares are reported both pooled across all segments and broken out per segment.
Group
A cluster of related query units, such as answers gathered under one prompt topic or category, used to check whether class shares hold steady across different subject matter rather than being an artifact of one topic.
Largest group share
The proportion of all units contributed by the single biggest group, used to confirm that no one query topic is large enough to distort the overall or per-engine class percentages.
Mitigation
A data-cleaning or deduplication step applied to the raw rows before classification, intended to prevent repeated or malformed rows from inflating a business's name count within a single answer.
Excluded rows
Rows removed from the dataset before analysis, for reasons such as malformed data. In this study that count is zero, meaning every collected row was classified and counted.

Questions about this study

So does Claude actually name more businesses per answer than ChatGPT, or not?
No, and this is the whole point of the study. When answers are sorted into count classes rather than judged by impression, Claude puts only 43.4% of its answers in the heaviest class, names 10 or more, while ChatGPT puts 49.7% of its answers there. Claude instead concentrates more in the lighter names 1 to 5 class, at 5.8% versus ChatGPT's 1.6%. The common assumption that Claude is the more list-heavy engine does not hold up once you count.
What counts as a 'names 10 or more' answer, exactly?
It is one of four classes used to bin every answer by how many distinct businesses it names inside a single response: names none, names 1 to 5, names 6 to 9, and names 10 or more. An answer goes into names 10 or more if it lists ten or more distinct businesses, regardless of how those businesses are formatted, ranked, or described. The classing is done per answer unit, not per query group, so an engine's class share is just the share of its own answers that fall into that bucket.
Where did this data come from and how much of it is there?
The study pooled 926,792 answer units across 193 query groups, collected between 2026-04-27 and 2026-09-18. Every unit was assigned to exactly one of four classes based on how many distinct businesses it named. No rows were excluded from the analysis, 0 were dropped, so the class shares reported describe the entire collected set, not a filtered subset of it.
Isn't it possible one engine just got easier or narrower prompts, which would explain the gap without saying anything about the engines themselves?
That is a real concern and the study checks for it directly. Across the 193 query groups, the share of answers in names 10 or more ranges from 5.7% to 73.6%, so topic and prompt phrasing clearly move the class an answer lands in. What the study cannot fully rule out is whether the three engines were tested against an identical distribution of query groups or merely a similar one. If prompt assignment was not matched group-for-group across engines, some of the ChatGPT-Claude gap could reflect which topics each engine happened to answer more of, rather than a stable behavioral difference. Matching or stratifying by query group would distinguish the two explanations.
Could one single giant query group be skewing the whole result?
Unlikely, based on the concentration numbers. The largest single query group accounts for only 1.5% of all units, and the average group holds 4,802.03 units across 193 groups. With no group anywhere near dominating the sample, it would take an unusual coincidence across many groups, not just one outlier, to manufacture the engine-level gap seen here.
What's the actual takeaway for someone building on top of Gemini instead of ChatGPT or Claude?
Gemini behaves differently from both. It has the highest names none share of the three engines, at 5.7% of its answers, meaning it is the most likely to name zero businesses at all, and the lowest names 10 or more share, at 24%. If your use case depends on getting a long list of named businesses back reliably, Gemini's answer distribution is the weakest fit of the three engines measured here, and you would want to test your specific prompts rather than assume any engine's overall tendency applies to your case.
Check our work

Data and method

The complete row-level dataset is published open and ungated under CC BY 4.0. Every number on this page can be recomputed from it.

Limitations we volunteer

  • Single pass. Run-to-run variance is not characterised.
  • Gemini's cited sources are largely unavailable through Google's API, so source analysis rests on the other engines.

How to cite this study

Counted properly, Claude names fewer businesses per answer than ChatGPT, not more. BusySeed, 2026-04-27. https://busyseed.com/research/distinct-businesses-per-question-anthropic-corrected

About BusySeed

BusySeed is a data-driven growth marketing agency that measures and improves how brands appear in AI-generated answers.

Free strategy session

Want to know how AI answers describe you?

We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.

Omar Jenblat, Founder & CEO of BusySeed
Omar JenblatFounder & CEO, BusySeed
  • Your category measured the same way
  • Your own numbers, not a sample deck
  • Fifteen minutes, no obligation

First, who are we meeting?

Three fields, then pick your time. We read up on you before the call so we open with something useful.

No sales sequence. If you never pick a time, we leave it there.