Does the Company AI Recommends First Stay the Same? A 15-Week Volatility Study
According to BusySeed, the company an AI assistant names first for a category changes from one week to the next in 56.4% of week-to-week transitions, and 82.3% of the question and engine pairs tracked saw their top answer change at least once.
According to BusySeed, the company an AI assistant names first for a category changes from one week to the next in 56.4% of week-to-week transitions, and 82.3% of the question and engine pairs tracked saw their top answer change at least once.
What to take away
- The identity of the top-mentioned company changed in 56.4% of consecutive week pairs across 480 tracked series, so a single snapshot rank is a weak guide to next week.
- 82.3% of series saw the leader change at least once over a median of 13 weeks, versus only 85 series where one company held the top spot the entire window.
- Gemini showed the highest week-to-week churn at 61.2% and ChatGPT the highest share of series with any change at 85.6%, meaning the three engines do not agree on how stable rankings are.
- Visibility itself is common, 95.9% of the 1664-company cohort appeared somewhere, but which single company appears first is the unstable part, not whether a company appears at all.
- Engines differ sharply in how often any given company shows up at all: 43.6% on ChatGPT, 42.1% on Gemini, and 29.3% on Claude, so a company's baseline exposure depends heavily on which engine is asked.
- None of the 1664 companies in the cohort were completely invisible across all three engines (0% zero-visibility rate), so the fight is over rank and first mention, not over presence versus absence.
- Because the dataset spans only 2 snapshot boundaries defining the window and one model version per engine, the volatility measured here reflects short-term answer variance, not necessarily long-run market share change.
Key findings
- 1 average mention position where they do appear, across 6385 appearances
- 43.6% of the cohort is visible on ChatGPT
- 29.3% of the cohort is visible on Claude
- 42.1% of the cohort is visible on Gemini
Why this question matters, and to whom
Every AI visibility product sells a number: a rank, a score, a position on a list. The implicit promise behind that number is that it describes something stable, a state of the world that a company can move up in, defend, and report to a board. If a vendor tells a marketing team "you are ranked third for this query on ChatGPT," the team reasonably assumes third place is a position, not a coin flip.
This study tests that assumption directly. It does not ask whether AI engines mention companies (they do, almost universally in this cohort, see the visibility findings page). It asks a narrower and more consequential question: if you ask the same engine the same commercial question again next week, does it name the same company first?
The answer matters to three groups differently. Buyers of AI visibility tracking need to know whether a single weekly read is a measurement or a sample from a noisy process, because that changes how they interpret a drop from first to third. Agencies pitching "AI SEO" services need to know whether a client's position can be durably won, or whether the same effort produces a different-looking result depending on which week the client happens to check. And engine builders and researchers care because leader instability at the level measured here, 56.4% of week-to-week comparisons showing a different top company, is itself a finding about how these models generate answers: it suggests the underlying retrieval and generation process is sensitive to small changes in phrasing, context, or model sampling rather than reading from a fixed, cached ranking.
The practical stakes are concrete. A company that pays for optimization work aimed at climbing to the top mention position needs to know how much of any subsequent change is signal and how much is the background rate of turnover that would have happened regardless. Without that baseline, every week's rank change looks like news. With it, most week-to-week movement looks like what this study found it to be: normal.
How the measurement works, and why it was built this way
The method is a repeated-measures design. The same 160 prompts were sent to each of three engines, ChatGPT, Claude, and Gemini, on a recurring weekly cadence across the run window from 2026-04-27 to 2026-08-03. Each prompt, on each engine, forms one time series; with 160 prompts and 3 engines, that produces 480 distinct series. Each series has a median of 13 weekly observations.
For each series, the study recorded which company appeared in the first mention position that week, using the same extraction method across all 59266 rows and 6385 usable answers in the underlying data. A "leader change" is defined as the top-mentioned company differing between one week's observation and the immediately following week's observation, within the same prompt and engine. Aggregating across all adjacent week pairs in all series gives 5905 total observed transitions, of which 3328 involved a different top company than the prior week, 56.4% of all transitions.
This design was chosen over a single before-and-after snapshot because a single pair of measurements cannot distinguish a real, durable shift from ordinary noise. A two-point comparison, "were you first in January and third in June," cannot tell you whether the change happened once and held or whether the top spot flickered every week in between and June happened to be a low week. Only a dense weekly series lets you calculate a turnover rate and separate persistent leaders (companies that hold the position across most or all weeks) from companies that occupy it for one week and are displaced the next.
Each engine was queried using a single model version across the entire window, recorded in the dataset as one entry per engine under model versions, so the volatility measured is not an artifact of the vendor silently upgrading the underlying model mid-study. That control matters: if Claude or Gemini had changed model versions partway through, any shift in leader identity could reflect a genuine model upgrade rather than ordinary run-to-run variance. Holding model version fixed isolates the latter.
The headline result, and what it does not mean
Across 480 series, the top-mentioned company changed between consecutive weekly observations 56.4% of the time. Put differently, if a company held first place this week, the unconditional chance it would still hold first place next week, averaged across the whole cohort, was 43.6%. That is barely better than even.
At the series level, the pattern is even starker. 82.3% of the 480 series experienced at least one leader change somewhere across their median 13 weeks of observation. Only 85 series, a small fraction of the total, kept a single company in first place for every observed week. A permanent, uncontested leader is the exception, not the rule.
What this does not mean is that the ranking process is random or meaningless. A 56.4% week-to-week change rate is consistent with several very different underlying processes, and the data as collected cannot fully separate them: (1) genuine and roughly continuous re-ranking as the engines' retrieval draws on updated or newly indexed content each week, (2) sampling variance in the language model's generation process, where the same prompt against the same underlying knowledge produces a different first-listed company simply because of how the model samples its output, and (3) sensitivity to trivial prompt-level differences, such as the automated system re-issuing a prompt with slightly different phrasing or context each week. Distinguishing these would require, at minimum, repeating the identical prompt multiple times within the same week to measure the pure sampling-variance component, then comparing that to the week-to-week component. That comparison is not part of this dataset and is a natural next study.
What the result does mean, without qualification, is that a single weekly snapshot of "who is ranked first" is not a stable descriptor of a company's standing. Any report, dashboard, or claim that treats one week's top mention as a fixed fact about the market is overstating the precision of the underlying measurement.
Is the change gradual or binary, and what follows from that
The shape of the instability matters for what a company should do about it. Two very different worlds produce a similar-looking aggregate turnover rate.
In the first world, change is gradual: a handful of companies trade the top spot back and forth in a tight, semi-predictable rotation, and the ecosystem has an effective "top tier" of three or four contenders who alternate. In this world, a company that is not in that top tier essentially never wins first place, no matter how much the wider field turns over.
In the second world, change is closer to binary and diffuse: many different companies take a turn in first place over the course of the window, with no small stable set dominating, and instead a wide field of eligible companies rotates through. In this world, more companies have a realistic shot at appearing first at some point, but no single win is durable.
The data collected here is most consistent with the second, more diffuse pattern. The gap between the share of series where the leader changed at least once, 82.3%, and the share of series with one permanent, unchanged leader, a small number of 85 out of 480, is wide enough that a narrow, stable top tier cannot explain most of the volatility. If a small set of three or four companies were simply trading places with each other, the study would still classify those series as "leader changed," but the number of distinct companies occupying first place across the full window would be small and repeating. Testing that distinction precisely, counting distinct leaders per series rather than just changes, is possible from the underlying data but is a finer cut than the headline figures reported here.
What follows practically: a company should not conclude that failing to rank first this week means it is locked out of a stable top tier it cannot break into. The more plausible reading is that first place is being distributed across a wider set of eligible companies over time, and that consistent visibility (appearing somewhere in the answer) matters more than any single week's top-position claim. That reframes the operational question from "how do we take and hold first place" to "how do we stay in the rotation."
What the engines are reading when they decide who to recommend
None of the three engines in this study expose their retrieval process directly, and this study did not capture the source citations the engines drew on when generating answers (the underlying dataset's fields for top-cited sources are empty here, an honest limitation rather than a null result). What can be said comes from the pattern of outputs, not from instrumentation of the engines' internals.
The basic architecture common to all three products is retrieval-augmented generation: the model does not have a single fixed lookup table of "the best company for X." Instead, each time it is asked, it draws on some combination of its training data and, for most production consumer AI products, a live or recent web retrieval step, then generates an answer conditioned on whatever it retrieved plus its own learned associations. Both halves of that process can vary run to run. The retrieval half varies because the underlying web content changes, indexes update, and the exact set of pages retrieved for a given prompt is not guaranteed to be identical across requests. The generation half varies because large language models sample from a probability distribution over next words rather than always outputting the single highest-probability continuation, so even an identical retrieval result can produce a different first-mentioned company from one run to the next.
This matters because the two sources of variance imply different remedies. If instability comes mainly from retrieval drift, a company's best lever is improving the durability and freshness of the content that gets retrieved: maintaining pages that are repeatedly re-indexed and cited, keeping claims current, and ensuring the company's own site and third-party mentions are consistently crawlable and current. If instability comes mainly from generation-side sampling, no amount of content work removes the variance, and the honest posture is to expect noise and measure it, not chase it.
The data in this study cannot cleanly separate these two sources, because it does not log the specific sources retrieved on each run. What it can say is that the appearance rate for commercial and transactional prompts is uniformly high across categories (100% appearance rate on 4355 transactional answers, 100% on 823 commercial-investigation answers), which suggests the engines reliably surface a plausible set of companies for these query types even while disagreeing on which one goes first.
Where the engines disagree, and why one score is misleading
A visibility product that reports a single blended score across engines hides a real and measurable disagreement between them. The three engines in this study do not behave the same way, on two separate dimensions: how often they change their mind, and how often they mention any given company at all.
Week-to-week leader change rate, by engine
| Engine | Series tracked | Leader change rate | Series with any change |
|---|---|---|---|
| ChatGPT | 160 | 58.9% | 85.6% |
| Claude | 160 | 49.1% | 80% |
| Gemini | 160 | 61.2% | 81.3% |
Gemini shows the highest week-to-week leader change rate at 61.2%, while Claude is the most stable at 49.1%. But ChatGPT has the highest share of series where the leader changed at least once across the window, 85.6%, meaning ChatGPT's instability is more widespread across prompts even though its per-week change rate is lower than Gemini's. These are not the same measurement and a single averaged number across the three would erase the difference.
The second disagreement is baseline visibility. On ChatGPT, 43.6% of tracked mentions (726 of the relevant base) showed a given company as visible; on Gemini it is 42.1% (700); on Claude it is markedly lower, 29.3% (488). Claude mentions a narrower set of companies than ChatGPT or Gemini do. A company that is well represented on ChatGPT and Gemini but absent from Claude's answers will see a single blended "AI visibility score" understate a real, structural gap on one specific engine, a gap that a business decision-maker (which engine's users matter most to this company) needs to see broken out, not averaged away.
The practical implication for anyone evaluating an AI visibility vendor: ask whether the reported score is a blend or is broken out by engine, and ask for the engine-level change rate, not just the current snapshot rank. A single number that folds together a 61.2% churn engine and a 49.1% churn engine describes neither well.
What a company on the wrong side of this should actually do
If first-place mention turns over on 56.4% of week-to-week comparisons, the first correction is to stop treating any single week's rank, high or low, as a verdict. A company that drops from first to third should ask whether that drop persisted across two or three subsequent weekly reads before treating it as a real decline, precisely because 82.3% of series experienced at least one change somewhere in the window without necessarily reflecting an underlying, sustained shift.
Second, the finding that only 0% of the 1664-company cohort was completely invisible across all three engines means the realistic goal for most companies is not "win first place" but "remain in the rotation." 95.9% of the cohort appeared somewhere across the three engines. The operational question is whether a company shows up reliably across weeks and across prompt types, not whether it captured the top slot in any single read.
Third, because the engines disagree with each other (Claude's baseline visibility rate of 29.3% versus ChatGPT's 43.6% and Gemini's 42.1%), a company should track engine-level results separately rather than a single blended score, and should identify which engine matters most for its actual customer base before investing effort. Optimizing generically for "AI visibility" without specifying the engine risks improving performance on the engine that already performed well while leaving a genuine gap on another engine untouched.
Fourth, prompt type matters. The appearance rate is high and roughly uniform across informational, navigational, commercial-investigation, and transactional prompts in aggregate, but that uniformity is measured across the whole cohort; a specific company's presence is likely concentrated in some prompt types and thin in others. A company should look at its own appearance pattern by prompt type (4355 transactional answers and 823 commercial-investigation answers make up the bulk of the volume) rather than assume uniform coverage.
Fifth, and most concretely, given that most week-to-week variance cannot currently be attributed cleanly to content freshness versus model sampling, the defensible response is to maintain current, accurate, and repeatedly crawlable content about the company's offering, since that is the lever within a company's control regardless of which source of variance dominates, and to measure exposure over a run of weeks rather than react to any single reading.
How to read the published dataset yourself
The dataset underlying this study contains 59266 rows and 6385 usable answers, drawn from 1664 companies across the categories listed in the by-category breakdown, covering a run window from 2026-04-27 to 2026-08-03 (2 boundary timestamps defining that window). Every row that entered the analysis carried an ok status, 59266 of 59266 rows, so the usable-answer count reflects rows that passed the study's basic quality checks, not the raw crawl total.
Each of the three engines, ChatGPT, Claude, and Gemini, was queried using a single, fixed model version across the entire window (one version per engine, recorded in the dataset's model version field), which is what allows the volatility findings to be attributed to run-to-run and week-to-week variance rather than to the vendor changing the underlying model mid-study.
A reader who wants to check a specific claim should look for four fields together: the leader-change count and the transitions-observed count for a given engine (their ratio is the leader change percentage reported here), the count of series where the leader ever changed against the total series count for that engine (that ratio is the "ever changed" percentage), the count of series with one permanent leader against the total series (a stricter, complementary check), and the visible-on-engine counts for the specific company or companies of interest. These four together let a reader verify whether a specific company's experience matches the cohort-wide pattern or diverges from it.
The study applied 3 mitigations to the raw data before analysis; a careful reader should ask a vendor what those mitigations were and whether they could plausibly inflate or deflate a turnover rate, since data cleaning choices (how ties in first position are broken, how a null or missing mention is treated, how multiple company name variants are reconciled) directly shape whether a given week-to-week comparison counts as a leader change. None of those specific mitigation steps are detailed in the summary tables here, which is itself worth noting as a limit on what can be independently verified from the published aggregates alone.
Findings in depth
Each of these has its own page, written to stand on its own.
If a company is recommended first this week, how likely is it to be recommended first again next week
How often does the top-mentioned company actually change from one week to the next
Across every tracked prompt and engine, the top mention turned over in 56.4% of consecutive weekly comparisons.
Read this finding →Out of all the prompt-engine combinations tracked, what share ever saw a different company take the top spot
How many tracked queries saw the leader change at least once over the study window
82.3% of 480 tracked series saw the top-mentioned company change at least once across a median of 13 weeks.
Read this finding →Is one AI engine more stable than the others, and does that matter for a blended visibility score
Do ChatGPT, Claude, and Gemini agree on how often the leader changes
Gemini's week-to-week leader change rate of 61.2% is well above Claude's 49.1%, meaning a single blended score across engines hides real differences.
Read this finding →If almost every company shows up somewhere in AI answers, why does the leader still change so often
Why near-universal visibility does not mean a stable number one
95.9% of the 1664-company cohort appeared somewhere across the three engines, yet only 85 of 480 tracked series kept one company in first place the whole time.
Read this finding →How this sits against other published work
Most published work on generative engines asks a different question than ours. Aggarwal et al.'s GEO: Generative Engine Optimization (arXiv, also published at KDD 2024 and archived on Princeton's Collaborate repository) treats a generative engine's output as something a publisher can act on: the paper introduces GEO-bench and shows that specific content changes can raise a source's visibility in generative answers by a stated margin. That paper assumes the target is stable enough to optimize toward. Our study does not test optimization at all; it tests stability itself, and we find the top-named company changes at least once in 82.3% of our 480 weekly series, with only 85 series holding a single leader the entire window. Neither result contradicts the other: a technique can reliably move a company up in a given week's answer while the identity of whoever sits in first place still turns over the following week for reasons unrelated to that technique, such as a change in which web pages the engine's retrieval step surfaced.
Where our design overlaps most closely with prior work is Semrush's three-month study of cited domains, which tracked weekly citation share across three LLMs and reported ChatGPT's citation rate for Reddit falling from close to 60% of prompt responses to around 10% within about six weeks, and Wikipedia's share falling from roughly 55% to less than 20% over the same study. That is citation-share volatility at the domain level, not leader-position volatility at the company level, but it is evidence of the same underlying phenomenon: what an engine surfaces on a repeated prompt is not fixed week to week. Our finding that week-to-week leader turnover runs at 56.4% across engines is consistent with, and extends, that pattern to a commercial-recommendation setting rather than a general citation-share setting.
Ahrefs' Brand Radar and Profound's citation-pattern studies take single-snapshot or short-window approaches, ranking domains like Reddit, Wikipedia, Amazon, Forbes, and Business Insider by citation share at one point in time, or comparing platforms' citation habits over a fixed quarter. These studies establish that engines differ from each other in what they tend to cite, which matches our own finding that leader-change rates differ by engine (58.9% for ChatGPT, 49.1% for Claude, 61.2% for Gemini). But none of them repeat the same prompts on a weekly cadence for a period as long as our 13-week median, so none can speak to whether a citation leader from one week predicts the next week's leader. That is the gap our design fills, at the cost of a narrower prompt set (160 prompts) than the hundreds of thousands of prompts Semrush or Profound draw on.
The vendor documentation from OpenAI and Anthropic explains a mechanism our data cannot: both companies confirm that citations are generated per-query from live web search rather than from a fixed, cached ranking, which is consistent with our result but does not by itself explain why turnover is as high as we observe. An alternative explanation we cannot rule out with this dataset is that our prompt wording itself, not the underlying commercial reality, is unstable in a way that invites different phrasing of the same answer each week; distinguishing that from genuine retrieval volatility would require holding prompts byte-for-byte identical and checking underlying search results directly, which our method does not do.
Finally, the buyer-behavior figures from Forrester (via Machine Relations) and G2 (via Demand Gen Report) establish why this instability matters commercially, buyers report using AI tools to compare and validate vendors before contact, and Gartner's reporting adds the caveat that most buyers still confirm AI output with a human seller. Our study measures the instability those figures presuppose is consequential; it does not measure whether buyers notice or act on a change in who is named first.
References
Sources this study reads against. Every link was fetched and confirmed reachable at publication.
- GEO: Generative Engine Optimization arXiv, 2023 Foundational paper formalizing generative engines and introducing GEO-bench; establishes that source-side optimization can shift visibility in generative engine answers, the premise our volatility study assumes but does not itself test.
- The Most-Cited Domains in AI: A 3-Month Study Semrush, 2026 Longitudinal weekly-citation study across three LLMs over thirteen weeks, the closest published precedent in method to our week-over-week design; documents large swings in citation share for individual domains over short periods.
- GEO: Generative Engine Optimization ACM SIGKDD (KDD 2024 Proceedings), 2024 Formal peer-reviewed conference record of the GEO paper, cited to confirm the work passed peer review rather than remaining a preprint.
- AI Visibility Statistics (2026): How Often Brands Appear in ChatGPT, Perplexity, Claude & Gemini Boring Marketing, 2026 The closest methodological precedent among industry blogs, running repeated brand-visibility audits across four engines; cited for its finding that engines differ substantially in how often they cite a given brand, and flagged as a non-peer-reviewed vendor source.
- ChatGPT search for Enterprise and Edu OpenAI Help Center, 2026 Companion vendor documentation describing the Sources panel, used to describe how citation and ranking are surfaced to end users.
- Gartner: AI Is Reshaping B2B Buying, but Human Sellers Still Close the Confidence Gap Demand Gen Report, 2026 Secondary report of a Gartner buyer survey; cited for the finding that most buyers still validate AI-generated vendor recommendations with a human seller, a caveat on how much a first-place mention actually decides.
- ChatGPT Search OpenAI Help Center, 2026 Vendor documentation confirming ChatGPT search responses can include inline citations, used to describe the mechanism by which a named company enters an answer.
- 100 Most Cited Domains in ChatGPT Ahrefs, 2026 Single-snapshot ranking of ChatGPT's most-cited domains via Ahrefs Brand Radar, used as a point-in-time contrast to our repeated-measures approach.
- AI Platform Citation Patterns: How ChatGPT, Google AI Overviews, and Perplexity Source Information Profound (tryprofound.com), 2026 Cross-engine comparison of citation concentration by source type, used to note that engines differ systematically in whom they cite, a factor relevant to why leader identity might vary by engine.
- 94% of B2B Buyers Use AI for Vendor Research (Forrester 2026 Buyers' Journey Survey, reported) Machine Relations, 2026 Secondary report of a Forrester buyer survey; cited for figures on how commonly AI tools are used to compare vendors before contact, establishing why first-place mentions matter commercially.
- How ChatGPT sources the web Profound (tryprofound.com), 2026 Large-sample study of ChatGPT citation behavior over a defined quarter, cited for its methodology description and for cross-platform citation comparisons.
- GEO: Generative Engine Optimization Princeton University, Collaborate repository, 2024 Institutional repository record confirming authorship and affiliation of the GEO paper.
- Web search tool Anthropic, Claude Platform Docs, 2026 Vendor documentation describing how Claude's web search tool retrieves and cites live web content, used to explain why Claude's answers can change week to week as underlying search results change.
- Half of B2B Software Buyers Now Start Their Research with AI Chatbots G2, via Demand Gen Report, 2026 Secondary report of a G2 buyer survey; cited for the finding that buyers think more highly of vendors an AI chatbot recommends, supporting the commercial stakes of leader-position volatility.
- Citations Anthropic, Claude Platform Docs, 2026 Documentation of how Claude chunks source text for citation, used to explain the mechanics of citation granularity referenced in our discussion of method limits.
Terms used in this study
- Leader
- The company named or listed first in an AI engine's answer to a given prompt at a given point in time. Being the leader is a position in a specific week's answer, not a permanent attribute of a company.
- Leader change
- An instance where the company in the leader position for a given prompt and engine differs between one weekly observation and the immediately following weekly observation.
- Series
- One tracked combination of a single prompt run against a single engine, observed repeatedly over time. This study tracked 480 series, one for each pairing of the 160 prompts with the 3 engines.
- Transition
- Any adjacent pair of weekly observations within a series, whether or not the leader actually changed between them. The count of transitions is the denominator used to calculate the leader change rate.
- Appearance rate
- The share of answers, within a given prompt type, in which a company (or, in the aggregate figures reported here, any eligible company) is mentioned at all, regardless of position.
- Visibility
- Whether a company is mentioned anywhere in an AI engine's answer to a relevant prompt, distinct from whether it is mentioned first. A company can be visible without ever being the leader.
- Retrieval-augmented generation
- The general architecture behind consumer AI chat products, in which the model draws on retrieved content, such as recent web pages, in addition to its own trained knowledge, before generating an answer. Both the retrieval step and the generation step can introduce variation between runs.
- Model version
- The specific release of an underlying AI model used to answer prompts. Holding this fixed across a study window ensures that measured volatility reflects run-to-run variance rather than the vendor silently upgrading the model mid-study.
- Run window
- The calendar period over which a study's repeated measurements were collected, given here as the two boundary dates 2026-04-27 and 2026-08-03.
- Cohort
- The full set of companies included in a study's measurement, here 1664 companies across the listed categories.
Questions about this study
What did this study actually measure?
So does the top-ranked company usually stay the same or not?
What counts as a 'leader change' here? Does a tiny wording shift count?
Couldn't this just be noise in how the AI phrases things, rather than real change in company standing?
Do all three engines agree on how volatile rankings are?
Is this the same thing as a company losing visibility or disappearing from results?
Data and method
The complete row-level dataset is published open and ungated under CC BY 4.0. Every number on this page can be recomputed from it.
Limitations we volunteer
- Single pass. Run-to-run variance is not characterised.
- Gemini's cited sources are largely unavailable through Google's API, so source analysis rests on the other engines.
How to cite this study
Does the Company AI Recommends First Stay the Same? A 15-Week Volatility Study. BusySeed, 2026-04-27. https://busyseed.com/research/category-leader-volatility-check
About BusySeed
BusySeed is a data-driven growth marketing agency that measures and improves how brands appear in AI-generated answers.
More BusySeed research
Do AI Engines Recommend the Same Companies? A 3-Engine Agreement Study
According to BusySeed, 88.4% of the company recommendations AI assistants make for a buyer question come from only one of three engines, and just 3.2% are made by all three.
AI Utah 100 AI-Visibility Study
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
