The three engines agree with each other on almost nothing.
Out of every pair of engine and company where at least one engine made a mention, the three engines landed on the same answer together only a small fraction of the time.
3.3% of 515,720 pairs
This study asked the same 22,875 prompts of three engines, ChatGPT, Claude, and Gemini, and recorded which companies each one named. Restricting attention to the 515,720 company-prompt pairs where at least one engine named a company at all, the three engines agreed with each other, meaning all three named the same company, only 3.3% of the time. Exactly two engines agreed 8.6% of the time. The remaining 88.1%, the large majority, were named by exactly one engine, with the other two either silent or naming someone else.
Why this denominator matters. These percentages are conditional on at least one engine having named a company for that prompt; they do not describe the full cohort, and they should not be read as "only 3.3% of companies are visible." A company entirely absent from all three engines' answers does not enter this table at all. What the table describes is the behavior of engines toward the subset of companies that at least one of them noticed.
Two explanations, and how to tell them apart. There are at least two plausible reasons agreement is this low. One is that the engines are making genuinely different judgments from similar information, weighting factors like recency, review sentiment, or category fit differently, and arriving at different shortlists as a result. The other is that the engines are working from different underlying information in the first place, different training data cutoffs, different live retrieval sources, different indexes, so that disagreement reflects incomplete overlap in what each engine knows rather than a disagreement in judgment applied to shared facts. This single month of data cannot distinguish the two. What would: tracking whether the same companies keep winning on the same engine month after month (consistent with a judgment difference) versus the winning engine rotating unpredictably across months (consistent with a noisier information-overlap story). That is the reason this study is designed to repeat monthly rather than run once.
What this means for buying a visibility score. Any product or report that claims a single "AI visibility ranking" without specifying which engine, or engines, it measured is describing, at best, one engine's behavior. Given that 88.1% of named pairs in this study came from a single engine acting alone, a one-engine score has a substantial chance of missing what a buyer sees on a different assistant entirely. The practical fix is not complicated: ask which engines were tested, and ask to see the agreement rate across them, the same way this page reports it, before trusting a single blended number.
What does not follow from this finding. It would be a mistake to conclude that AI visibility is therefore unmeasurable or arbitrary. The disagreement rate is itself a stable, repeatable measurement, comparable month over month, and a company's position within it, which engine names it and which does not, is itself useful information, just not information that collapses cleanly into one number.
Other findings
What share of companies actually get mentioned by AI assistants at all?
Almost every company shows up somewhere. That is not the same as showing up where it counts.
Nearly all of the cohort was named by at least one engine during the study window, but that headline figure hides how unevenly that presence is distributed.
Does it matter which AI engine a buyer happens to use?
Gemini names companies less often than ChatGPT or Claude do.
Across the same set of prompts, the three engines named a cohort company at noticeably different rates, with Gemini the least likely of the three to name anyone from the cohort.
Once a company is mentioned by an AI assistant, where does it typically land in the list?
Being named is not the same as being named first.
The average position of a mention across all three engines sits in the middle of a typical answer, and the average shifts slightly depending on which engine is asked.
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
- Your category measured the same way
- Your own numbers, not a sample deck
- Fifteen minutes, no obligation
