Do ChatGPT, Claude, and Gemini agree on how often the leader changes

Gemini's week-to-week leader change rate of 61.2% is well above Claude's 49.1%, meaning a single blended score across engines hides real differences.

Is one AI engine more stable than the others, and does that matter for a blended visibility scoreMeasured 2026-04-27
The number

61.2% vs 49.1%

Anyone comparing AI engines on a single combined visibility score is averaging over meaningfully different behavior. This study tracked leader volatility separately for each of the three engines, ChatGPT, Claude, and Gemini, each with 160 tracked series (the same count for all three, since each of the 160 prompts was run on all three engines).

The three engines by the numbers

EngineWeek-to-week leader change rateShare of series with any changeTotal leader changes observed
ChatGPT58.9%85.6%1162
Claude49.1%80%971
Gemini61.2%81.3%1195

Gemini has the highest per-week churn, changing its top-mentioned company on 61.2% of adjacent week comparisons, against 1974 transitions observed for ChatGPT and 1977 for Claude, out of similar overall transition counts per engine. Claude is the most stable on this specific measure, at 49.1%.

But stability at the week-to-week level is not the same as stability across the whole window. ChatGPT actually has the highest share of series that saw a change at some point in the study, 85.6%, even though its per-week churn rate is lower than Gemini's. That combination, lower per-week churn but a higher share of series eventually affected, suggests ChatGPT's changes may arrive in occasional, more decisive shifts rather than continuous flickering, while Gemini's higher per-week rate suggests more continuous back-and-forth. This study's aggregate figures cannot fully confirm that distinction; testing it would require looking at where in each series the changes cluster, not just how many occurred.

Why this matters for a blended score

A visibility product that reports one number across all engines necessarily averages Gemini's 61.2% against Claude's 49.1%, landing near the overall 56.4% figure. That blended number describes no single engine well. A company whose customers primarily interact with Claude-based products would be reading a volatility estimate roughly 56.4% when the figure that actually applies to it is closer to 49.1%, a meaningfully more stable picture.

The same logic applies to baseline visibility, not just volatility. Claude's overall visibility rate in this cohort, 29.3%, is well below ChatGPT's 43.6% and Gemini's 42.1%. A company that is well represented on ChatGPT and Gemini but weak on Claude will have that weakness diluted into invisibility by a single blended score.

What to ask a vendor

Before relying on any single "AI visibility score," a buyer should ask whether the figure is broken out per engine, and specifically request the per-engine leader change rate and per-engine visibility rate, since a single blended number by construction cannot distinguish a company doing well on one engine and poorly on another from a company performing evenly across all three.

Free strategy session

Want to know how AI answers describe you?

We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.

Get my visibility read