Questions about this study
Do AI Engines Recommend the Same Companies? A 3-Engine Agreement Study
14 questions answeredMethod challenges included
What did this study actually measure?
It measured whether three AI engines, ChatGPT, Claude, and Gemini, name the same companies when given identical prompts. Across 2,500 prompts run on all three engines, the study logged 170,457 answer rows, kept 170,457 after quality checks, and identified a cohort of 30,564 distinct companies that showed up in at least one answer. For every company-prompt-engine combination, the study recorded whether the company was mentioned and, if so, at what position in the list. The question is not which companies are best, but how much the three engines agree with each other about which companies to mention at all.
So do the engines agree with each other or not?
Mostly not. Of 56,819 company-prompt pairs where at least one engine named a company, 88.4% were named by only one engine, and only 3.2% were named by all three. The remaining 8.4% landed in the middle, named by exactly two engines. That middle group matters because it shows agreement is not a coin flip between total overlap and total disagreement, it is a gradient, with most pairs sitting at the low end.
Does that mean most companies are invisible to two out of three engines?
Not invisible, just not consistently named. No company in the 30,564-company cohort failed to appear on any engine at all (0% never appeared anywhere). But 96% of the cohort appeared on at least one engine without appearing on all three. So the typical company is seen by some engines and missed by others on any given prompt, which is different from being permanently excluded by any one engine.
Why would three engines answering the same prompt give different lists?
The engines differ in training data, retrieval methods, and whatever ranking heuristics sit behind their answer generation, and none of that is visible from the outside in this study. A plausible mechanical explanation is that each engine draws from a different slice of the web or weights sources differently, so the same query surfaces a different shortlist. An alternative explanation is that the underlying company set is genuinely large and any single engine only samples a fraction of it per answer. This study cannot distinguish those two causes, it can only show that the resulting lists overlap less than someone might expect.
Is this just because one engine writes shorter answers than the others?
That is a real possibility this study cannot rule out. If one engine tends to list fewer companies per answer, it would show lower visibility rates without necessarily disagreeing on which companies matter. What the study does show is that ChatGPT surfaced 42.7% of companies it was asked about, Claude 39.9%, and Gemini 33.6%, a real spread. But answer length per engine was not separately measured here, so the gap could reflect list length, source selection, or both, and this data cannot separate them.
When a company does get named by more than one engine, does it show up in the same spot?
Yes, roughly. Average mention position is close across engines, Gemini at 4.85, ChatGPT at 5.25, and Claude at 5.38, against an overall average of 5.18 across 65,241 mentions. That means the disagreement between engines is mostly about whether a company gets named, not about where it lands once it is. A company that appears tends to land in a similar rank whichever engine produced the answer.
What counts as a 'usable' answer, and what got thrown out?
The study logged 170,457 raw answer rows and kept 170,457 after quality checks, meaning every logged row passed and none were discarded for errors (170,457 marked ok). That's a clean run in this instance, but it also means the quality bar itself is not visible in this dataset; readers cannot tell from these numbers alone what would have disqualified a row, only that nothing was disqualified this time.
How many companies and prompts are we talking about, concretely?
The study ran 2,500 distinct prompts against each of the 3 engines, drawing on a pool of 30,564 ranked companies and surfacing a working cohort of 30,564 companies that actually appeared in at least one answer. Categories ranged widely, from Blockchain Testing Services to Landscaping Services, with category sizes varying from small (nine companies in Healthcare IT Services) to large (nine hundred ninety four in SaaS). The prompt types included commercial-investigation, informational, navigational, and transactional queries.
Were there any engine versions or settings that might explain the gap?
Each engine was run under a single model version during the study window: one version each for Claude, Gemini, and ChatGPT, all logged as 'rankxa-warehouse' in the run metadata. The measurement window itself spanned two timestamps in 2026. Because only one version per engine was tested, this study cannot say whether a different model version, a different temperature setting, or a different prompt phrasing would close or widen the agreement gap. It describes one snapshot of one configuration per engine, not a general property of each engine's architecture.
If I only track one AI engine for my company's visibility, what am I missing?
You would systematically miss a large share of mentions. Since 88.4% of company-prompt pairs were named by only one engine, a monitoring setup built on a single engine misses most of the mentions that exist across the other two. A company could be doing well on Claude and invisible on ChatGPT in the same week, and a single-engine dashboard would never show that gap. The practical fix is to pull mentions from all three engines separately rather than building one composite score from one source.
Does 'never appears' mean some companies are just bad at this and get filtered out entirely?
No, the opposite: 0% of companies in the cohort never appeared on any engine, meaning zero companies were completely absent everywhere. Every company in the 30,564-company cohort was named by at least one engine at least once. The unevenness is about which engine names a given company on a given prompt, not about some companies being permanently excluded from all AI answers.
What does 'position' mean in this context, and why does it matter?
Position refers to where in an engine's answer a company is mentioned, first, third, tenth, and so on, on the assumption that earlier mentions are more prominent to a reader. The study found average position hovering near 5.18 across 65,241 mentions, with little spread between engines. This matters because it isolates two separate questions, whether a company is mentioned and where, and shows the disagreement between engines lives almost entirely in the first question, not the second.
Could this whole result just be noise from having too few prompts?
It is a fair challenge, and the honest answer is that 2,500 prompts across 3 engines produced 56,819 company-prompt pairs, which is a large base for the overall agreement figures. But the per-category breakdowns include some very small groups, such as nine companies in Healthcare IT Services or twenty four in Fintech Testing Services, and conclusions drawn at that category level rest on much thinner evidence than the headline numbers. Anyone citing a category-specific agreement rate should check the category's company count first.
What should a company actually do with this finding?
Stop treating any single AI engine's answer as a proxy for how visible you are everywhere. Since agreement across all three engines happened in only 3.2% of pairs, a company that wants to know its real AI visibility needs to query ChatGPT, Claude, and Gemini separately and track them as three distinct channels, similar to how a marketer would not judge search visibility from one search engine alone. The concrete Monday-morning change is adding the other two engines to whatever monitoring already exists for the first one.
Free strategy session
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
Omar JenblatFounder & CEO, BusySeed
- Your category measured the same way
- Your own numbers, not a sample deck
- Fifteen minutes, no obligation
