Questions about this study

Does the Company AI Recommends First Stay the Same? A 15-Week Volatility Study

13 questions answeredMethod challenges included
What did this study actually measure?
It measured whether the company an AI engine names first, when asked the same commercial question repeatedly, stays the same from one week to the next. The study ran three engines, ChatGPT, Claude, and Gemini, against 160 prompts on a recurring basis over a run window from 2026-04-27 to 2026-08-03, producing 480 separate weekly series (one per prompt per engine) with a median of 13 weeks of observations each. For each series it tracked which company held the top-mentioned spot and whether that identity changed between consecutive weeks.
So does the top-ranked company usually stay the same or not?
Mostly not. Across all tracked series, the company in first place changed in 56.4% of consecutive week pairs, and 82.3% of series saw the leader change at least once over the study window. Only 85 of 480 series kept one company in first place the entire time. A single week's top mention is a weak predictor of the following week's top mention on any of the three engines tested.
What counts as a 'leader change' here? Does a tiny wording shift count?
A leader change means the company occupying the number-one mention position in a weekly answer differed from the company occupying that position the previous week, for the same prompt and engine. It does not measure how far a company moved in the ranking or whether it dropped out entirely, only whether the specific company named first was different. A company could stay visible and near the top and still register as a change simply by trading places with another company.
Couldn't this just be noise in how the AI phrases things, rather than real change in company standing?
That is a real possibility and the study cannot fully rule it out. Each engine here has one model version (1 version tracked per engine), and the run window covers 2 snapshot boundaries, so what looks like volatility could partly reflect answer-generation variance rather than any underlying shift in the companies themselves. Distinguishing the two would require repeating the identical prompt multiple times within the same week to see how much churn exists with no time passing at all, which this design does not do.
Do all three engines agree on how volatile rankings are?
No. Gemini showed the highest week-to-week churn, with the identity of the top company changing in 61.2% of consecutive week pairs, while Claude was lowest at 49.1%. ChatGPT had the highest share of series where the leader changed at least once, 85.6%. Since the engines disagree on both the rate and the breadth of change, a company's experience of stability depends heavily on which engine someone is asking.
Is this the same thing as a company losing visibility or disappearing from results?
No, and this is an important distinction. None of the 1664 companies in the cohort had zero visibility across all three engines (0% zero-visibility rate), and 95.9% of the cohort appeared somewhere. What is unstable is which single company sits in first place, not whether companies show up in answers at all. A company can be consistently mentioned and still lose and regain the top spot repeatedly.
Why does visibility differ so much between engines?
On ChatGPT, 43.6% of the cohort appeared at least once; on Gemini it was 42.1%; on Claude it was 29.3%. The study did not investigate why, but plausible mechanisms include differences in each model's training data, how each engine selects and truncates the number of companies it names in a given answer, and how each handles ties or borderline relevance. Whatever the cause, a company's baseline exposure is not portable across engines, being visible on one says nothing definite about visibility on another.
What should a company actually do with this finding?
Treat a single week's first-place mention as a snapshot, not a status. Because the leader changed in 56.4% of consecutive week pairs, a company that lands in the top mention one week should expect a real chance of being displaced the next, and one that is displaced should not assume the loss is permanent. The practical move is to track mention position over several weeks per engine before concluding that a competitor has taken over or that a campaign improved standing, and to check standing separately on each engine rather than assuming one result generalizes.
Are there companies that held the top spot the whole time, and who are they?
The study counted 85 series, out of 480 total, where a single company held first place for every observed week. The underlying dataset does not report a public list of which companies these were per series, so this figure describes how common permanent leadership is, not which specific companies achieved it. Separately, the five companies with the broadest overall visibility in the dataset were Fiverr, HubSpot, Hibu, Big Wheel, and Scorpion, though high visibility does not necessarily mean holding first place continuously.
What is a 'series' in this study, exactly?
A series is one specific prompt asked to one specific engine, tracked repeatedly over the run window. Because there were 160 distinct prompts and 3 engines, the study produced 480 series in total. Each series has its own week-by-week record of which company appeared first, and the median series has 13 weeks of observations, though the number of weeks per series can vary.
Does this study tell us anything about long-term market share or real business performance?
No, and the study is explicit about that limit. It covers a run window bounded by 2 snapshot dates, from 2026-04-27 to 2026-08-03, with one model version tracked per engine. That is enough to describe short-term answer variance within this window but not enough to say whether any company's underlying market position is rising or falling. A company could show high week-to-week volatility in AI mentions while its actual market share is stable, or vice versa, and this dataset cannot distinguish those cases.
How many companies and data points does this study cover in total?
The dataset includes 59266 total rows and 6385 usable answers, covering 1664 companies across 1664 ranked companies overall. Of those rows, 59266 returned an ok status, meaning the query completed and produced a usable answer. This scale is what allows the study to report percentages like the leader-change rate with reasonable stability, though it does not change the fact that the window itself is short.
If ChatGPT has the highest visibility rate, does that mean it's the best engine to target for exposure?
It means ChatGPT showed the highest share of the cohort appearing at least once, 43.6%, compared to 42.1% on Gemini and 29.3% on Claude, within this specific set of prompts and window. But ChatGPT also had the highest share of series where the top spot changed at least once, 85.6%, so higher visibility there does not mean higher stability once a company reaches first place. Whether to prioritize an engine depends on whether the goal is broad mention or a defensible top position, and this study suggests those are different targets.
Free strategy session

Want to know how AI answers describe you?

We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.

Get my visibility read