Methodology
Do AI Engines Recommend the Same Companies? A 3-Engine Agreement Study
Engines and exact versions
A study that does not name the model version it ran against is not reproducible, because the answer changes when the model does.
| Engine | Model version |
|---|---|
| Claude | rankxa-warehouse |
| Gemini | rankxa-warehouse |
| ChatGPT | rankxa-warehouse |
Run window, UTC: 2026-05-23T06:48:44Z to 2026-08-01T22:00:00Z.
Method
The design starts from a fixed prompt set, 2,500 distinct prompts covering commercial-investigation, informational, navigational, and transactional intents, spanning categories such as SaaS, Legal Services, Real Estate Investment, and Manufacturing Services. Each prompt was sent to all three engines, ChatGPT, Claude, and Gemini, using a single model version per engine (1 version tracked for ChatGPT, 1 for Claude, 1 for Gemini), so differences in results reflect differences between engines rather than between model generations within an engine.
Runs were collected across a measurement window bounded by two timestamps in 2026, and 3 mitigations were applied during collection to reduce known sources of noise, such as retrying failed calls and normalizing company names before matching. The raw log contains 170,457 rows; after filtering to rows with a clean, parseable answer, 170,457 remained, all logged with an 'ok' status (170,457 of them), meaning none were dropped for parsing failure in the final count.
For every usable answer, the study recorded which companies were named and, when named, the ordinal position of the mention within the answer (first company mentioned, second, and so on). This produced two separate measurements that are easy to conflate but answer different questions. The first is whether a company appears at all on a given engine for a given prompt, which is the basis for the agreement analysis. The second is where in the answer it appears when it does, which is the basis for position analysis such as 5.18 as the overall average position across 65,241 scored mentions.
The unit of comparison for agreement is the company-prompt pair: for each prompt, which companies did each engine name, and how many of the three engines named each one. This is deliberately stricter than asking 'did the engines cover similar topics,' because it requires exact company identity to match, not just category or theme. A pair counts as agreed only if the same company name, after normalization, was produced by more than one engine for the same prompt.
Limitations we volunteer
Written by us, before anyone else found them.
- Single pass. Run-to-run variance is not characterised.
- Gemini's cited sources are largely unavailable through Google's API, so source analysis rests on the other engines.
Terms used in this study
- Company-prompt pair
- One company and one prompt considered together. If a company is named by at least one engine in response to a given prompt, that combination counts as one pair. Agreement is measured by how many of the three engines named that same pair.
- Appearance rate
- The share of answers, within a given group such as a prompt type, in which at least one tracked company was named by the engine.
- Visibility rate
- The share of companies in the cohort that were named by a specific engine at least once across the study's prompts.
- Mention position
- The ordinal place a company holds within an engine's answer when it lists multiple companies, for example first, second, or third mentioned. A lower average position means companies tend to be named earlier in the answer.
- Sole mention
- A company-prompt pair named by exactly one of the three engines and not named by the other two for that same prompt.
- Cohort
- The full set of distinct companies that appeared at least once, on any engine, anywhere in the study's usable answers.
- Usable answer
- A logged engine response that passed quality checks, meaning it could be parsed cleanly enough to extract company names and positions, and was marked with an 'ok' status.
- Prompt intent type
- One of four categories, commercial-investigation, informational, navigational, or transactional, describing the kind of question a prompt represents, based on how a real buyer might phrase it.
- Category
- A business vertical or service type used to group companies and prompts, such as Blockchain Testing Services or Landscaping Services.
- Visible somewhere, not everywhere
- A company that was named by at least one engine but not by all three, the group most affected by which single engine a buyer happens to consult.
References
Sources this study reads against. Every link was fetched and confirmed reachable at publication.
- GEO: Generative Engine Optimization arXiv, 2023 Introduces GEO-bench and reports that generative engine optimization techniques can raise a source's visibility in generative engine answers by up to 40%, establishing that citation behavior in these systems is measurable and manipulable.
- AI Platform Citation Source Index 2026 5WPR via PR Newswire, 2026 Industry citation-source ranking across five engines, cited for its figure on Reddit's dominance as a cited source and its scale of citations analyzed.
- GEO: Generative Engine Optimization Princeton University, Collaborate repository, 2023 Institutional archival record of the GEO paper, used to confirm authorship and publication context.
- AI Visibility Statistics (2026) boringmarketing.com, 2026 Independent industry dataset reporting brand citation rates by engine, cited for comparison against our per-engine visibility rates.
- New G2 Research: Half of B2B Software Buyers Now Start Their Research With AI Chatbots G2 via PR Newswire, 2026 Survey evidence that buyers act on AI chatbot recommendations, establishing why engine disagreement in company naming has commercial consequence.
- GEO: Generative Engine Optimization ACM SIGKDD (Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining), 2024 Peer-reviewed proceedings version of the GEO paper, cited as the field's foundational benchmark for measuring visibility in generative engine responses.
- Tow Center's Latest Report on AI Search Engines Columbia Journalism School, 2025 Institutional news page confirming the scope and release of the Tow Center report.
- Generative engine optimization Wikipedia, 2026 Background and definitional context for the term GEO as used in this related-work discussion.
- AI search engines fail to produce accurate citations in over 60% of tests, according to new Tow Center study Nieman Lab, 2025 Summary of the Tow Center study reporting per-engine failure rates, including Perplexity's lowest failure rate among the eight engines tested.
- Gartner Survey Finds 69% of B2B Buyers Turn to Sales Reps to Validate AI-Generated Insights Gartner, 2026 Survey data on how buyers source and validate vendor information, including average number of information sources used, cited as context for why cross-engine agreement matters to buyers.
- GEO: Generative Engine Optimization OpenReview, 2023 Review version of the GEO paper, cited for its framing that generative engines need a purpose-built visibility metric distinct from search-engine ranking metrics.
- AI Search Has a Citation Problem Columbia Journalism Review / Tow Center for Digital Journalism, 2025 Primary independent study testing eight AI search engines on citation accuracy for news content, the closest existing benchmark to our cross-engine comparison.
Data
The complete row-level dataset is published open and ungated under CC BY 4.0. Every number in this study can be recomputed from it.
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
- Your category measured the same way
- Your own numbers, not a sample deck
- Fifteen minutes, no obligation
