Questions about this study
After correcting for Anthropic's duplicate rows, how many distinct businesses does Claude actually name per buyer question, and does it differ from OpenAI or Gemini
10 questions answeredMethod challenges included
So does Claude actually name more businesses per answer than ChatGPT, or not?
No, and this is the whole point of the study. When answers are sorted into count classes rather than judged by impression, Claude puts only 43.4% of its answers in the heaviest class, names 10 or more, while ChatGPT puts 49.7% of its answers there. Claude instead concentrates more in the lighter names 1 to 5 class, at 5.8% versus ChatGPT's 1.6%. The common assumption that Claude is the more list-heavy engine does not hold up once you count.
What counts as a 'names 10 or more' answer, exactly?
It is one of four classes used to bin every answer by how many distinct businesses it names inside a single response: names none, names 1 to 5, names 6 to 9, and names 10 or more. An answer goes into names 10 or more if it lists ten or more distinct businesses, regardless of how those businesses are formatted, ranked, or described. The classing is done per answer unit, not per query group, so an engine's class share is just the share of its own answers that fall into that bucket.
Where did this data come from and how much of it is there?
The study pooled 926,792 answer units across 193 query groups, collected between 2026-04-27 and 2026-09-18. Every unit was assigned to exactly one of four classes based on how many distinct businesses it named. No rows were excluded from the analysis, 0 were dropped, so the class shares reported describe the entire collected set, not a filtered subset of it.
Isn't it possible one engine just got easier or narrower prompts, which would explain the gap without saying anything about the engines themselves?
That is a real concern and the study checks for it directly. Across the 193 query groups, the share of answers in names 10 or more ranges from 5.7% to 73.6%, so topic and prompt phrasing clearly move the class an answer lands in. What the study cannot fully rule out is whether the three engines were tested against an identical distribution of query groups or merely a similar one. If prompt assignment was not matched group-for-group across engines, some of the ChatGPT-Claude gap could reflect which topics each engine happened to answer more of, rather than a stable behavioral difference. Matching or stratifying by query group would distinguish the two explanations.
Could one single giant query group be skewing the whole result?
Unlikely, based on the concentration numbers. The largest single query group accounts for only 1.5% of all units, and the average group holds 4,802.03 units across 193 groups. With no group anywhere near dominating the sample, it would take an unusual coincidence across many groups, not just one outlier, to manufacture the engine-level gap seen here.
What's the actual takeaway for someone building on top of Gemini instead of ChatGPT or Claude?
Gemini behaves differently from both. It has the highest names none share of the three engines, at 5.7% of its answers, meaning it is the most likely to name zero businesses at all, and the lowest names 10 or more share, at 24%. If your use case depends on getting a long list of named businesses back reliably, Gemini's answer distribution is the weakest fit of the three engines measured here, and you would want to test your specific prompts rather than assume any engine's overall tendency applies to your case.
What does 'measured properly' even mean here, versus however this was measured before?
It means every answer was placed into one of four fixed count classes, names none, names 1 to 5, names 6 to 9, or names 10 or more, based on a straight count of distinct businesses named, and then class shares were compared engine by engine across the full 926,792-unit dataset. The informal alternative this study argues against is judging an engine's list-heaviness from a handful of read transcripts, which can easily be swayed by which examples happen to get read or remembered. Counting every unit and binning consistently removes that selection effect, though it depends on the counting rule itself being applied the same way across engines, which is a mitigation the study reports handling.
Is the modal answer really just one or none names, or does it usually name a decent number?
Across the pooled set of 926,792 rows, the single largest class is names 6 to 9, at 51.8% of all answers, and names 10 or more adds another 39.4%. Answers naming zero businesses, names none, are the rarest class at 2.3%. So the typical answer across engines lists a substantial handful of businesses, six or more names is the norm, not the exception.
What data cleaning was done before these percentages were calculated, and should I trust it?
Three mitigations were applied to the data before classing rows, and the process excluded 0 rows, meaning nothing was dropped from the collected set. The study states that three cleaning steps were applied but does not itemize what each one corrected, which is worth knowing if you plan to reproduce the classing yourself: you can verify the class shares against the full 926,792-unit set, but you cannot independently check what the mitigations changed without more detail on each step.
If I'm choosing between ChatGPT and Claude for a use case that needs long lists of named businesses, what should I do with this finding?
Lean toward ChatGPT if your prompts resemble the ones in this study, since 49.7% of its answers landed in the names 10 or more class against 43.4% for Claude. But treat this as a starting point, not a guarantee, because the class an answer falls into varies enormously by query group, from 5.7% to 73.6% across the 193 groups measured. The safest move is to run your own actual prompts against both engines and count the results yourself rather than assuming the aggregate gap transfers to your specific topic.
Free strategy session
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
Omar JenblatFounder & CEO, BusySeed
- Your category measured the same way
- Your own numbers, not a sample deck
- Fifteen minutes, no obligation
