How often do ChatGPT, Claude, and Gemini name the same company for the same prompt

Full three-way agreement is rare; most company mentions come from just one engine.

When you ask ChatGPT, Claude, and Gemini the same question, how often do they name the same companyMeasured 2026-05-23
The number

88.4% of pairs

If you have ever asked two different AI assistants the same question about which company to hire or buy from, and gotten two different shortlists, this finding explains how often that happens across a large sample rather than anecdotally.

The study sent 2,500 prompts to three engines, ChatGPT, Claude, and Gemini, and recorded which companies each engine named in its answer. For every prompt, a company-prompt pair exists for each company named by at least one engine on that prompt. Across the dataset, 56,819 such pairs were named by at least one engine.

Of those, 88.4% (50,209 pairs) were named by exactly one engine. 8.4% (4,798 pairs) were named by exactly two engines. Only 3.2% (1,812 pairs) were named by all three engines.

In plain terms: for every roughly nine pairs where a company gets named by at least one engine, about eight of those are cases where only one engine mentioned that company at all for that prompt. Three-way consensus, where every engine independently arrived at the same company for the same question, happened on only about three pairs out of every hundred.

Why would this be true rather than false? Three plausible mechanisms, not mutually exclusive. First, each engine may draw on a differently weighted set of training and retrieval sources, so one engine's shortlist candidates are not fully available to another. Second, each engine applies its own internal ranking or selection logic when compressing 'the market' down to a short answer, and with many plausible candidates in a category, small differences in that logic produce different final picks even from similar underlying knowledge. Third, answer length and style differ by engine, and an engine that tends to name more companies per answer will mechanically show up in more pairs, which inflates apparent coverage without necessarily reflecting better knowledge.

What would make this finding false, or at least much weaker? If the three-way agreement rate were closer to the sole-mention rate, engines would be behaving more like independent measurements of the same underlying reality, the way three analysts reading the same annual report might converge on the same top vendors. That is not what the data show. The gap between 88.4% and 3.2% is large enough that convergence is the exception, not the rule.

A reasonable objection: maybe the engines agree on category-level judgments, such as 'this is a strong vendor,' even when they disagree on exact company identity, perhaps due to name variants or close competitors being treated as interchangeable. This study cannot rule that out; it matched on normalized company identity, not on semantic similarity of company profiles, so near-agreement in spirit but not in exact naming would show up as disagreement here. That is a real limit of a name-matching method, and it means the true rate of 'engines converge on similar-caliber companies' could be somewhat higher than the rate of 'engines converge on the identical company,' which is what this finding measures.

For a company, the immediate implication is that being checked on one engine tells you little about your standing on another. For a buyer of a monitoring tool, it means a single-engine visibility score understates how much of the competitive landscape it is missing, since 88.4% of the mentions in this dataset exist on only one engine.

Free strategy session

Want to know how AI answers describe you?

We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.

Get my visibility read