The three engines agree with each other on almost nothing.

Out of every pair of engine and company where at least one engine made a mention, the three engines landed on the same answer together only a small fraction of the time.

If ChatGPT recommends a company, does Claude or Gemini recommend the same one?Measured 2026-07-01
The number

3.3% of 515,720 pairs

This study asked the same 22,875 prompts of three engines, ChatGPT, Claude, and Gemini, and recorded which companies each one named. Restricting attention to the 515,720 company-prompt pairs where at least one engine named a company at all, the three engines agreed with each other, meaning all three named the same company, only 3.3% of the time. Exactly two engines agreed 8.6% of the time. The remaining 88.1%, the large majority, were named by exactly one engine, with the other two either silent or naming someone else.

Why this denominator matters. These percentages are conditional on at least one engine having named a company for that prompt; they do not describe the full cohort, and they should not be read as "only 3.3% of companies are visible." A company entirely absent from all three engines' answers does not enter this table at all. What the table describes is the behavior of engines toward the subset of companies that at least one of them noticed.

Two explanations, and how to tell them apart. There are at least two plausible reasons agreement is this low. One is that the engines are making genuinely different judgments from similar information, weighting factors like recency, review sentiment, or category fit differently, and arriving at different shortlists as a result. The other is that the engines are working from different underlying information in the first place, different training data cutoffs, different live retrieval sources, different indexes, so that disagreement reflects incomplete overlap in what each engine knows rather than a disagreement in judgment applied to shared facts. This single month of data cannot distinguish the two. What would: tracking whether the same companies keep winning on the same engine month after month (consistent with a judgment difference) versus the winning engine rotating unpredictably across months (consistent with a noisier information-overlap story). That is the reason this study is designed to repeat monthly rather than run once.

What this means for buying a visibility score. Any product or report that claims a single "AI visibility ranking" without specifying which engine, or engines, it measured is describing, at best, one engine's behavior. Given that 88.1% of named pairs in this study came from a single engine acting alone, a one-engine score has a substantial chance of missing what a buyer sees on a different assistant entirely. The practical fix is not complicated: ask which engines were tested, and ask to see the agreement rate across them, the same way this page reports it, before trusting a single blended number.

What does not follow from this finding. It would be a mistake to conclude that AI visibility is therefore unmeasurable or arbitrary. The disagreement rate is itself a stable, repeatable measurement, comparable month over month, and a company's position within it, which engine names it and which does not, is itself useful information, just not information that collapses cleanly into one number.

Free strategy session

Want to know how AI answers describe you?

We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.

Omar Jenblat, Founder & CEO of BusySeed
Omar JenblatFounder & CEO, BusySeed
  • Your category measured the same way
  • Your own numbers, not a sample deck
  • Fifteen minutes, no obligation

First, who are we meeting?

Three fields, then pick your time. We read up on you before the call so we open with something useful.

No sales sequence. If you never pick a time, we leave it there.