Questions about this study
Do AI Engines Recommend the Same Companies? A 3-Engine Agreement Study
12 questions answeredMethod challenges included
What does it mean that 88.4 percent of recommendations came from only one engine?
When a company appeared in an AI engine's answer to a buyer question, there was an 88.4 percent chance that neither of the other two engines mentioned that same company for the same question. Out of 56819 company-question pairs where at least one engine named a company, 50209 were sole mentions. This means each engine is largely drawing from its own internal model of what companies are relevant, not converging on a shared set of obvious answers. If a buyer asks the same question to ChatGPT versus Claude versus Gemini, they will usually get different company recommendations.
How often do all three AI engines agree on recommending the same company?
All three engines agreed on the same company for the same question just 3.2 percent of the time. Of 56819 total company-question pairs, only 1812 achieved unanimous recommendation. This is a very low consensus rate and suggests that each engine has distinct criteria, training data, or retrieval mechanisms shaping its answers. Unanimous agreement appears to be the exception rather than the rule.
Which AI engine recommends the most companies?
ChatGPT recommended the largest share of the cohort, appearing with 42.7 percent of the 30564 companies in the study. Claude followed at 39.9 percent, and Gemini recommended the fewest at 33.6 percent. In absolute terms, ChatGPT named 13057 companies, Claude named 12206, and Gemini named 10284. The gap between the highest and lowest reach engines is roughly nine percentage points, which means a company's measured AI visibility can vary substantially depending on which engine you check.
Could the low agreement just be because you asked vague questions?
The study used 2500 prompts spanning four intent types: transactional, navigational, informational, and commercial investigation. Agreement rates were consistent across all types, ranging from 36.9 percent appearance rate for navigational queries to 38.7 percent for informational ones. If vague prompts were the cause, we would expect tighter agreement on more specific, transactional prompts, but that pattern did not emerge. The consistency suggests the divergence reflects genuine differences in how each engine models relevance, not prompt ambiguity.
What should I do if my company shows up on one engine but not the others?
This is the situation for 96 percent of companies in the study: visible somewhere but not everywhere. The practical implication is that your AI visibility strategy cannot rely on a single engine as a proxy for the whole market. First, measure your presence on each engine separately. Second, investigate why you appear where you do. The engines may weight different signals, such as structured data, citation patterns, or entity associations. Third, treat the engine where you are absent as a separate optimization problem rather than assuming gains on one will transfer.
Why would the same company get recommended by one engine and ignored by another?
Several mechanisms could explain this. Each engine has different training data cutoffs, so one may have indexed more recent information about a company. They also differ in how they retrieve and rank entities, with some weighting citations or structured knowledge graphs more heavily. Additionally, the engines may have different thresholds for confidence before naming a company. A company that barely clears the relevance bar for ChatGPT might fall just below it for Gemini. The study cannot isolate which factor dominates, but the 88.4 percent sole-mention rate suggests these internal differences are substantial.
Were any companies completely invisible to all three engines?
No. The study found 0 companies that appeared in zero engines, meaning 0 percent of the 30564 cohort was never mentioned. Every company in the study appeared in at least one engine's responses at some point across the 2500 prompts. However, this does not mean every company has strong visibility. A company could appear once for an obscure query and remain effectively invisible for the questions that matter to its buyers.
Does the type of question affect how much the engines agree?
Surprisingly little. Transactional prompts, which are typically purchase-oriented, had an appearance rate of 38.3 percent. Navigational prompts, where a user is looking for a specific company or site, had 36.9 percent. Informational prompts came in at 38.7 percent, and commercial investigation at 38.6 percent. The narrow range suggests that engine disagreement is a structural feature, not something that varies by query intent. Buyers asking any kind of question should expect to see different recommendations depending on which engine they use.
How large was this study and how was it conducted?
The study queried three AI engines, ChatGPT, Claude, and Gemini, with 2500 distinct prompts and recorded 170457 recommendation events involving 30564 unique companies. Data collection ran between May and August 2026. Each prompt was sent to all three engines, and the responses were parsed to identify which companies were named. The analysis then measured how often engines agreed or disagreed on specific company-question pairs. The large sample, 56819 pairs where at least one engine named a company, provides confidence that the low agreement rate is not a small-sample artifact.
What is a 'company-question pair' and why does it matter?
A company-question pair is one instance where a specific company was mentioned in response to a specific prompt. The study counted 56819 such pairs across all engines. This unit of analysis matters because it lets us ask: when engine A recommends company X for question Y, do engines B and C also recommend X for Y? Measuring agreement at the pair level reveals how consistent the engines are on the same exact task, rather than just comparing their overall company coverage. The finding that 88.4 percent of pairs are sole mentions shows that engines rarely converge on the same answer to the same question.
If I only track my visibility on one engine, how wrong could my picture be?
Potentially very wrong. ChatGPT shows 42.7 percent of companies while Gemini shows 33.6 percent, a gap of roughly nine percentage points. More importantly, 96 percent of companies that appear somewhere do not appear everywhere. If you measure only on the engine where you happen to perform well, you may miss blind spots on engines your buyers also use. A single-engine visibility score reflects that engine's opinion, not a consensus reality. Multi-engine measurement is necessary for an accurate picture.
What are the limitations of this study?
The study measures which companies engines name, not whether those recommendations lead to conversions or are accurate. It also cannot see inside the engines to explain why they diverge. The 2500 prompts, while numerous, may not represent every query a real buyer would ask. Additionally, AI engines update their models over time, so agreement rates measured in mid-2026 may shift as training data and retrieval systems evolve. Finally, the study treats each engine as a single entity, but users may experience variation based on conversation history or other personalization factors not captured here.
Free strategy session
Get my visibility read
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
