Do AI Engines Recommend the Same Companies? A 3-Engine Agreement Study
According to BusySeed, 88.4% of the company recommendations AI assistants make for a buyer question come from only one of three engines, and just 3.2% are made by all three.
According to BusySeed, 88.4% of the company recommendations AI assistants make for a buyer question come from only one of three engines, and just 3.2% are made by all three.
What to take away
- When one engine recommends a company for a buyer question, the other two engines do not recommend it 88.4 percent of the time.
- Only 3.2 percent of company-question pairs achieved unanimous recommendation from all three engines.
- A single-engine AI visibility score reflects that engine's opinion, not a consensus reality.
- ChatGPT names the most companies (42.7 percent of the cohort), Gemini the fewest (33.6 percent), creating a 42.7-to-33.6 percent gap depending on where you measure.
- Agreement rates hold steady across prompt types, from transactional (38.3 percent) to navigational (36.9 percent).
- 96 percent of companies visible somewhere are not visible everywhere, meaning nearly all companies have engine-specific blind spots.
Key findings
- 88.4% of company recommendations were made by only one of the 3 engines
- 3.2% were made by all 3 engines for the same question
- 5.18 average mention position where they do appear, across 65241 appearances
- 42.7% of the cohort is visible on ChatGPT
- 39.9% of the cohort is visible on Claude
- 33.6% of the cohort is visible on Gemini
Why Engine Agreement Matters for AI Visibility
When a buyer asks an AI assistant for recommendations, they ask one assistant, not three. They might use ChatGPT because it is built into their workflow, Claude because their company has an enterprise contract, or Gemini because they are already in Google's ecosystem. The buyer does not run a consensus poll.
This creates a measurement problem for anyone tracking "AI visibility." If the three major engines recommend substantially different companies for the same question, then any score derived from a single engine describes that engine's opinion, not a stable property of the brand. A company might score well on one platform and poorly on another, and neither score would be wrong, just incomplete.
The stakes are practical. Marketing teams are beginning to treat AI recommendations the way they once treated search rankings: as a channel to monitor and optimize. Agencies sell AI visibility audits. Brands ask whether they are "showing up" in AI answers. But if the three engines disagree more often than they agree, then the answer to "are we showing up?" is always "it depends on where you look."
This study does not attempt to measure which engine gives better recommendations. That would require knowing the buyer's true preferences, which the data does not contain. Instead, it measures something narrower and checkable: when one engine names a company, how often do the others name it too?
Who this matters to
- Brands tracking AI visibility: A single-engine score may not generalize. A company visible on ChatGPT is not necessarily visible on Claude or Gemini.
- Agencies selling visibility services: Promising to improve "AI visibility" without specifying which engine is an underspecified promise.
- Buyers using AI for recommendations: The company you get depends on the assistant you ask. This is not a flaw in the assistants; it is a property of how recommendation engines work when they are trained on overlapping but not identical data.
- Researchers studying AI recommendations: Agreement rates provide a baseline for understanding how much engine choice matters versus question specifics.
How the Measurement Works
The unit of analysis is a (question, company, engine) triple. For each of the 2500 prompts in the sample, the study collected responses from all 3 engines. Every company named by at least one engine for a given question enters the denominator for that question.
The agreement calculation
For each company-question pair, the study counts how many engines named that company:
| Engines naming the company | Pairs | Percent of total |
|---|---|---|
| 1 (sole mention) | 50209 | 88.4 |
| 2 (partial agreement) | 4798 | 8.4 |
| 3 (unanimous) | 1812 | 3.2 |
The denominator is 56819 pairs where at least one engine made a recommendation.
Design choices and their consequences
Deterministic sampling. Questions were selected by ordering prompt IDs by their MD5 hash, which produces a reproducible sample without storing a random seed. The same database will always yield the same sample.
Most recent run only. Some questions were run multiple times. Only the most recent completed run per (question, engine) was used, so a question re-run ten times contributes one observation, not ten.
Company universe per question. The denominator is not "all companies that exist" but "companies an engine was willing to recommend for this question." A company that no engine ever mentions for a given question does not appear in that question's denominator. This means the study measures disagreement among active recommendations, not the gap between recommendations and reality.
What this cannot see
The method cannot determine which engine is right. If ChatGPT recommends Company X and Claude does not, Company X might be a good recommendation Claude missed, or a bad recommendation ChatGPT should not have made. Agreement is not accuracy. High agreement could mean all three engines share the same bias.
The Headline Result: Engines Mostly Disagree
Of 56819 company-question pairs where at least one engine made a recommendation, 50209 (88.4 percent) were named by only one engine. Just 1812 pairs (3.2 percent) achieved unanimous agreement.
This is a statement about disagreement, not error. When ChatGPT recommends a company that Claude and Gemini do not mention, three explanations are possible:
- ChatGPT found a good company the others missed. Perhaps its training data includes sources the others lack, or its retrieval system surfaced a relevant but obscure provider.
- ChatGPT made a questionable recommendation. Perhaps it hallucinated a company name, or recommended a company that is no longer in business, or picked up on a marketing-heavy source that overstated the company's relevance.
- The question has many valid answers and the engines are sampling from the space of good options. For a query like "best digital marketing agencies," there may be dozens of reasonable recommendations, and each engine picks a different subset.
The data cannot distinguish these cases. What it can establish is that relying on a single engine's recommendations, or measuring visibility on a single engine, captures less than half the picture.
The practical implication
If you measure your company's AI visibility on ChatGPT alone, you are measuring ChatGPT's opinion. That opinion may differ substantially from Claude's or Gemini's. 96 percent of companies that appear somewhere do not appear everywhere. This is not a rounding error; it is the dominant pattern.
A company that appears in all three engines for a question is rare enough to be notable. A company that appears in only one engine for a question is the norm.
The Shape of Disagreement: Binary, Not Gradual
The distribution of agreement is heavily skewed. Nearly nine in ten company-question pairs (88.4 percent) are sole mentions. Partial agreement, where exactly two engines name the same company, accounts for just 8.4 percent of pairs. Unanimous agreement is rarer still at 3.2 percent.
This distribution suggests that engines are not pulling from a shared list and occasionally disagreeing about edge cases. Instead, they appear to be constructing substantially different recommendation sets, with occasional overlap.
Engine-level differences
The engines do not recommend the same number of companies overall:
| Engine | Companies ever mentioned | Percent of cohort |
|---|---|---|
| ChatGPT | 13057 | 42.7 |
| Claude | 12206 | 39.9 |
| Gemini | 10284 | 33.6 |
ChatGPT names the most companies, Gemini the fewest. The gap between them, roughly nine percentage points, is large enough that a company visible on the most expansive engine might be invisible on the most selective one.
Position when mentioned
When engines do mention companies, they tend to mention them at similar positions:
| Engine | Average mention position |
|---|---|
| ChatGPT | 5.25 |
| Claude | 5.38 |
| Gemini | 4.85 |
| Overall | 5.18 |
Gemini mentions companies slightly earlier on average (position 4.85), while Claude mentions them slightly later (position 5.38). The differences are small enough that position alone does not explain the disagreement. The engines are not agreeing on who to recommend and merely disagreeing on order; they are disagreeing on who to recommend at all.
Agreement Rates by Prompt Type
One might expect different levels of agreement for different kinds of questions. A navigational query ("How do I contact Acme Corp?") seems more constrained than a transactional query ("Who should I hire for mobile app development?"). Perhaps engines agree more on questions with fewer valid answers.
The data does not support this hypothesis. Appearance rates across prompt types are nearly identical:
| Prompt type | Total appearances | Appearance rate |
|---|---|---|
| Transactional | 47628 | 38.3 percent |
| Commercial investigation | 8457 | 38.6 percent |
| Informational | 4469 | 38.7 percent |
| Navigational | 4687 | 36.9 percent |
The range spans just 38.6 percent to 36.9 percent, a gap of less than two percentage points. Navigational queries, which one might expect to have clearer "right answers," show the lowest appearance rate.
Why this might be
Several factors could explain the consistency:
- All prompt types in this sample ask for company recommendations. Even a navigational query in this context involves naming a company. The engines face similar tasks regardless of prompt framing.
- The sample may not include truly constrained queries. A query like "What is IBM's phone number?" has one right answer, but such queries may not appear in a sample designed to study company recommendations.
- Engine disagreement may be driven more by training data differences than query characteristics. If the engines have access to different information about companies, that gap persists regardless of how the question is phrased.
The takeaway for practitioners: do not assume that clearer queries will produce more consistent recommendations across engines. The 88.4 percent sole-mention rate appears to be structural, not dependent on query type.
What Engines Are Reading When They Decide
The study does not have direct visibility into each engine's retrieval process, but the pattern of agreement suggests some hypotheses about what drives recommendations.
Shared information should produce shared outputs
If all three engines had access to identical information about every company, and used similar reasoning to evaluate that information, they would produce similar recommendations. The 88.4 percent sole-mention rate implies that at least one of these conditions fails.
Training data differences. Each engine is trained on a different corpus. A company with strong coverage in sources favored by ChatGPT's training may appear in ChatGPT's recommendations but not in Claude's or Gemini's. This is not a matter of one engine being better; it is a matter of different inputs producing different outputs.
Retrieval system differences. Modern AI assistants often combine base model knowledge with real-time retrieval. The retrieval systems may prioritize different sources or apply different relevance criteria.
Response generation differences. Even given identical retrieved information, the models may weight evidence differently. One might prioritize recency, another might prioritize source authority, a third might prioritize keyword matches.
What a company can influence
A company cannot control which sources each engine trusts. But it can increase its presence across the sources that engines collectively draw from:
- Authoritative directories that appear in multiple engines' training and retrieval pipelines
- Industry publications that engines treat as credible for the company's category
- Structured data (schema markup, knowledge graph entries) that helps engines verify company information
- Consistent entity information across the web, reducing the chance that an engine sees conflicting signals
The goal is not to optimize for one engine at the expense of others, but to build a presence that survives across different retrieval and ranking systems.
The Problem with Single-Engine Visibility Scores
Many AI visibility tools and audits measure presence on a single engine, often ChatGPT due to its market share. This study quantifies why that approach is incomplete.
The math of single-engine measurement
ChatGPT mentions 42.7 percent of the cohort. Claude mentions 39.9 percent. Gemini mentions 33.6 percent. But these are not subsets of each other. A company visible on ChatGPT is not necessarily visible on Claude.
29329 companies (96 percent of those visible anywhere) are visible on at least one engine but not all three. If you measure only on ChatGPT, you miss the companies that appear only on Claude or Gemini. More importantly, you may tell a client they are "visible" when they are visible on one engine and invisible on two.
What this means for visibility audits
A responsible AI visibility audit should:
- Specify which engines were measured. "You are visible" should always be followed by "on ChatGPT, Claude, and/or Gemini."
- Report per-engine results separately. An aggregate score hides whether a company has broad coverage or narrow.
- Acknowledge what was not measured. Bing's Copilot, Perplexity, and other AI assistants may produce different results.
- Update periodically. Engine behavior changes as models are updated and retrieval systems are refined.
What this means for brands
If a competitor is visible on all three engines and you are visible on one, you are at a disadvantage that a single-engine score would not reveal. Conversely, if you are invisible on the engine your best customers happen to use, your aggregate score may overstate your actual visibility to buyers.
The 88.4 percent sole-mention rate means that engine choice is not a minor detail. It is often the determining factor in whether a company appears at all.
What a Company on the Wrong Side Should Do
A company that appears on one engine but not the others faces a specific problem: whatever signals earned that single recommendation are not transferring to the other engines. This section outlines a diagnostic approach.
Step 1: Identify the gap
Before optimizing, understand the current state:
- Which engines mention your company, for which questions?
- Where in the response does your company appear?
- Are competitors mentioned that you are not?
This study provides data for 30564 companies across 2500 prompts. For companies in the dataset, the raw data shows which engines named them and when.
Step 2: Audit your information footprint
Engines draw from different sources. Audit your company's presence on:
- Major business directories (industry-specific and general)
- Review platforms relevant to your category
- Professional associations that maintain member directories
- Wikipedia and Wikidata (for larger companies)
- News coverage in outlets engines may cite
If you are strong in sources that ChatGPT favors but weak in sources Claude or Gemini favor, that gap may explain the disagreement.
Step 3: Fix entity consistency
Engines struggle when they see conflicting information. Common problems:
- Different company names on different platforms (Inc. vs. LLC vs. no suffix)
- Outdated addresses or phone numbers
- Inconsistent descriptions of services offered
- Multiple websites for the same company
Consistency helps engines verify that mentions of your company across sources refer to the same entity.
Step 4: Monitor, do not assume
Engine behavior changes. A company invisible today may appear after a model update. A company visible today may disappear. Build ongoing monitoring into your process, preferably across all major engines rather than the one you find most convenient.
The goal is not to game any single engine, but to be the kind of company that multiple systems, with different data and different reasoning, independently conclude is worth recommending.
How to Read the Published Dataset
The study covers 30564 companies across 2500 prompts, generating 170457 usable recommendation events. The data is structured to support both company-level and question-level analysis.
Key columns
- prompt_id: Unique identifier for the question. The same prompt_id across engines means the same question was asked.
- company: The company name as returned by the engine.
- engine: Which engine (ChatGPT, Claude, or Gemini) made this recommendation.
- position: Where in the response the company was mentioned (lower is earlier).
- prompt_type: The query classification (transactional, informational, commercial-investigation, or navigational).
Common analyses
Company-level visibility: Filter to a specific company and count how many prompts each engine recommended it for. Compare across engines to see where the company is strong or weak.
Question-level agreement: For a specific prompt_id, list all companies mentioned and by which engines. Calculate the agreement rate for that question.
Category analysis: Filter to prompts in a specific industry category. The data covers categories from Accessibility Testing Services (392 companies) to Web Development Services (506 companies).
Limitations of the dataset
- Snapshot in time. Data was collected between May and August 2026. Engine behavior may have changed since.
- No ground truth. The data shows what engines said, not whether they were right.
- Prompt sample, not exhaustive. 2500 prompts is a sample, not the universe of possible buyer questions.
- Company names as returned. The same company may appear with slightly different names across engines; normalization was applied but may not be perfect.
Reproducing the sample
The sample was selected by ordering prompt IDs by their MD5 hash. Any researcher with access to the same prompt database can reproduce the sample without needing a stored random seed.
Findings in depth
Each of these has its own page, written to stand on its own.
How often does only one AI engine recommend a company when other engines do not?
The Sole-Mention Rate: Why Most AI Recommendations Are Unique to One Engine
When an AI engine recommends a company for a buyer question, that recommendation is unique to that engine 88.4 percent of the time. The other two engines simply do not mention the same company.
Read this finding →How often do all three AI engines agree on recommending the same company?
Only 3.2 Percent of Recommendations Are Unanimous
Unanimous agreement across ChatGPT, Claude, and Gemini is rare. When all three engines were asked the same buyer question, they all named the same company just 3.2 percent of the time.
Read this finding →Which AI engine recommends the most companies?
ChatGPT Recommends More Companies Than Claude or Gemini
ChatGPT mentioned 42.7 percent of the companies studied, compared to 39.9 percent for Claude and 33.6 percent for Gemini. The engines have different thresholds for recommendation.
Read this finding →Do AI engines agree more on certain types of buyer questions?
Query Type Does Not Predict Engine Agreement
Whether a prompt is transactional, informational, or navigational, engines agree at roughly the same rate. Appearance rates range only from 36.9 to 38.6 percent.
Read this finding →How this sits against other published work
Related Work
Research on generative engine optimization and AI citation behavior has grown rapidly since 2024, but most published work examines single engines or measures domain-level citation patterns rather than company-level recommendation agreement across multiple platforms. Our study addresses a gap that existing work identifies but does not fill.
Cross-Engine Divergence in Citation Behavior
The clearest precedent for expecting disagreement comes from Peec AI's analysis of 30 million sources across ChatGPT, Gemini, Perplexity, and Google AI Mode. That study found platform-level divergence in domain preferences: Google's AI Mode and AI Overviews favor social content, while ChatGPT leans toward editorial sources. Wikipedia, for instance, appears prominently in ChatGPT and Perplexity citations but is nearly absent from Google's platforms. Our finding that 88.4 percent of company-question pairs were named by only one engine is consistent with this pattern, extending it from domain-level divergence to company-level recommendation disagreement. The Peec AI study, however, does not measure agreement rates directly; it reports which domains each engine prefers, not how often engines converge on the same entity for the same query.
Citation Concentration Within Engines
Search Engine Land's coverage of Kevin Indig's research found that roughly 30 domains capture 67 percent of citations within a topic in ChatGPT. This concentration suggests that even within a single engine, recommendation behavior is skewed toward a small set of sources. Our study measures a related but distinct phenomenon: whether that concentration is consistent across engines. The answer appears to be no. With only 3.2 percent of pairs achieving unanimous recommendation from all three engines, the concentrated sets each engine favors do not substantially overlap at the company level.
Academic Foundations
The foundational academic work on generative engine optimization is Aggarwal et al.'s GEO paper, published at KDD 2024. That paper introduced GEO-bench, a benchmark for measuring visibility in generative engine responses, and demonstrated that optimization strategies vary in efficacy across domains. This domain-specificity finding supports our observation that engine behavior is not uniform, though the GEO paper focuses on how content producers can improve their visibility rather than on measuring cross-engine agreement. Importantly, GEO studied single-engine responses; our three-engine design directly tests whether gains on one platform translate to others.
Commercial Relevance
The practical stakes of cross-engine disagreement depend on how buyers actually use these tools. Grey Matter's synthesis of Forrester and Gartner data reports that 45 percent of B2B buyers used generative AI during a recent purchase. If a company is visible on ChatGPT (42.7 percent of our cohort) but not on Gemini (33.6 percent), the buyer's engine choice determines whether that company enters consideration. This asymmetry is invisible to any single-engine measurement and has direct implications for resource allocation.
Accuracy Concerns
The Columbia Journalism Review's Tow Center audit found that no publisher was spared inaccurate representations of its content in ChatGPT, regardless of affiliation with OpenAI. Our study does not assess accuracy; we measure only whether a company is named. A company could be mentioned by all three engines and still be misrepresented by all three. This limitation means our visibility metrics are necessary but not sufficient for understanding a company's AI presence.
What Our Method Adds
Prior work measures which domains or entities appear in AI responses. Our contribution is measuring agreement: how often the same company appears for the same prompt across engines. With 56819 total pairs where at least one engine named a company, we can compute agreement rates that domain-level studies cannot. The trade-off is that we do not analyze citation sources or assess factual accuracy; we treat each engine's output as given.
References
Sources this study reads against. Every link was fetched and confirmed reachable at publication.
- ChatGPT citations favor a small group of domains: Study Search Engine Land, 2025 Reports Kevin Indig's finding that roughly 30 domains capture 67 percent of citations within a topic in ChatGPT, establishing the concentration pattern that single-engine measurement would miss.
- The Most-Cited Domains in AI: A 3-Month Study Semrush, 2025 Cross-platform study finding Reddit and LinkedIn among the top five most-cited domains on ChatGPT, Google AI Mode, and Perplexity, suggesting social content has growing weight in AI responses.
- GEO: Generative Engine Optimization ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2024 Foundational academic work introducing GEO-bench for measuring visibility in generative engine responses; demonstrates that optimization strategies vary in efficacy across domains, supporting the premise that engine behavior differs by category.
- Top domains cited by AI search: Analysis based on 30M sources Peec AI, 2025 Large-scale analysis of 30 million sources across ChatGPT, Gemini, Perplexity, and Google AI Mode finding that platforms diverge in domain preferences, with Wikipedia strong for ChatGPT but absent from Google platforms.
- How B2B Buyers Use AI to Choose Vendors Grey Matter, 2026 Synthesizes Forrester and Gartner surveys showing 45 percent of B2B buyers used generative AI in a recent purchase, establishing commercial relevance of engine-level visibility differences.
- ChatGPT Search OpenAI Help Center, 2025 Primary vendor documentation explaining that ChatGPT responses using search include inline citations users can hover over and click to view sources.
- 100 Most Cited Domains in ChatGPT Ahrefs, 2025 Continuously updated dataset tracking domains cited in ChatGPT responses across US queries; provides single-engine benchmark against which multi-engine disagreement can be compared.
- How ChatGPT Search (Mis)represents Publisher Content Columbia Journalism Review (Tow Center), 2025 Independent audit documenting inaccuracies in ChatGPT citations regardless of publisher affiliation with OpenAI, underscoring that visibility alone does not guarantee accurate representation.
- Grounding with Google Search Google AI for Developers, 2025 Primary vendor documentation explaining how Gemini generates inline citations via groundingMetadata, clarifying the technical mechanism underlying Gemini's citation behavior.
- Web search tool Anthropic (Claude Platform Docs), 2025 Primary vendor documentation describing how Claude's web search generates targeted queries and returns citations to source materials.
Terms used in this study
- Sole mention
- A company-question pair where only one of the three engines named the company. The other two engines gave a response but did not include that company.
- Unanimous agreement
- A company-question pair where all three engines (ChatGPT, Claude, and Gemini) independently named the same company in their responses.
- Company-question pair
- The combination of a specific buyer question (prompt) and a specific company. The study counts how many engines recommended each pair.
- Appearance rate
- The frequency at which companies appear in engine responses, calculated as appearances divided by total possible slots.
- Prompt type
- A classification of buyer intent. Transactional queries seek to complete an action, informational queries seek knowledge, navigational queries seek a specific destination, and commercial investigation queries compare options before purchase.
- Engine reach
- The share of total companies that an engine ever mentions across all questions in the sample. Higher reach means the engine recommends more companies.
- Deterministic sample
- A sample selected by a repeatable process (here, MD5 hash ordering) rather than a random seed. Any researcher with the same database can reproduce the same sample.
- Position
- Where in an engine's response a company was mentioned. Position 1 means the company was the first named; higher numbers mean it appeared later in the response.
Questions about this study
What does it mean that 88.4 percent of recommendations came from only one engine?
How often do all three AI engines agree on recommending the same company?
Which AI engine recommends the most companies?
Could the low agreement just be because you asked vague questions?
What should I do if my company shows up on one engine but not the others?
Why would the same company get recommended by one engine and ignored by another?
Data and method
The complete row-level dataset is published open and ungated under CC BY 4.0. Every number on this page can be recomputed from it.
Limitations we volunteer
- Single pass. Run-to-run variance is not characterised.
- Gemini's cited sources are largely unavailable through Google's API, so source analysis rests on the other engines.
How to cite this study
Do AI Engines Recommend the Same Companies? A 3-Engine Agreement Study. BusySeed, 2026-05-23. https://busyseed.com/research/cross-engine-disagreement-buyer-questions
About BusySeed
BusySeed is a data-driven growth marketing agency that measures and improves how brands appear in AI-generated answers.
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
