Methodology
Does the Company AI Recommends First Stay the Same? A 15-Week Volatility Study
Engines and exact versions
A study that does not name the model version it ran against is not reproducible, because the answer changes when the model does.
| Engine | Model version |
|---|---|
| Claude | rankxa-warehouse |
| Gemini | rankxa-warehouse |
| ChatGPT | rankxa-warehouse |
Run window, UTC: 2026-04-27 to 2026-08-03.
Method
The method is a repeated-measures design. The same 160 prompts were sent to each of three engines, ChatGPT, Claude, and Gemini, on a recurring weekly cadence across the run window from 2026-04-27 to 2026-08-03. Each prompt, on each engine, forms one time series; with 160 prompts and 3 engines, that produces 480 distinct series. Each series has a median of 13 weekly observations.
For each series, the study recorded which company appeared in the first mention position that week, using the same extraction method across all 59266 rows and 6385 usable answers in the underlying data. A "leader change" is defined as the top-mentioned company differing between one week's observation and the immediately following week's observation, within the same prompt and engine. Aggregating across all adjacent week pairs in all series gives 5905 total observed transitions, of which 3328 involved a different top company than the prior week, 56.4% of all transitions.
This design was chosen over a single before-and-after snapshot because a single pair of measurements cannot distinguish a real, durable shift from ordinary noise. A two-point comparison, "were you first in January and third in June," cannot tell you whether the change happened once and held or whether the top spot flickered every week in between and June happened to be a low week. Only a dense weekly series lets you calculate a turnover rate and separate persistent leaders (companies that hold the position across most or all weeks) from companies that occupy it for one week and are displaced the next.
Each engine was queried using a single model version across the entire window, recorded in the dataset as one entry per engine under model versions, so the volatility measured is not an artifact of the vendor silently upgrading the underlying model mid-study. That control matters: if Claude or Gemini had changed model versions partway through, any shift in leader identity could reflect a genuine model upgrade rather than ordinary run-to-run variance. Holding model version fixed isolates the latter.
Limitations we volunteer
Written by us, before anyone else found them.
- Single pass. Run-to-run variance is not characterised.
- Gemini's cited sources are largely unavailable through Google's API, so source analysis rests on the other engines.
Terms used in this study
- Leader
- The company named or listed first in an AI engine's answer to a given prompt at a given point in time. Being the leader is a position in a specific week's answer, not a permanent attribute of a company.
- Leader change
- An instance where the company in the leader position for a given prompt and engine differs between one weekly observation and the immediately following weekly observation.
- Series
- One tracked combination of a single prompt run against a single engine, observed repeatedly over time. This study tracked 480 series, one for each pairing of the 160 prompts with the 3 engines.
- Transition
- Any adjacent pair of weekly observations within a series, whether or not the leader actually changed between them. The count of transitions is the denominator used to calculate the leader change rate.
- Appearance rate
- The share of answers, within a given prompt type, in which a company (or, in the aggregate figures reported here, any eligible company) is mentioned at all, regardless of position.
- Visibility
- Whether a company is mentioned anywhere in an AI engine's answer to a relevant prompt, distinct from whether it is mentioned first. A company can be visible without ever being the leader.
- Retrieval-augmented generation
- The general architecture behind consumer AI chat products, in which the model draws on retrieved content, such as recent web pages, in addition to its own trained knowledge, before generating an answer. Both the retrieval step and the generation step can introduce variation between runs.
- Model version
- The specific release of an underlying AI model used to answer prompts. Holding this fixed across a study window ensures that measured volatility reflects run-to-run variance rather than the vendor silently upgrading the model mid-study.
- Run window
- The calendar period over which a study's repeated measurements were collected, given here as the two boundary dates 2026-04-27 and 2026-08-03.
- Cohort
- The full set of companies included in a study's measurement, here 1664 companies across the listed categories.
References
Sources this study reads against. Every link was fetched and confirmed reachable at publication.
- GEO: Generative Engine Optimization arXiv, 2023 Foundational paper formalizing generative engines and introducing GEO-bench; establishes that source-side optimization can shift visibility in generative engine answers, the premise our volatility study assumes but does not itself test.
- The Most-Cited Domains in AI: A 3-Month Study Semrush, 2026 Longitudinal weekly-citation study across three LLMs over thirteen weeks, the closest published precedent in method to our week-over-week design; documents large swings in citation share for individual domains over short periods.
- GEO: Generative Engine Optimization ACM SIGKDD (KDD 2024 Proceedings), 2024 Formal peer-reviewed conference record of the GEO paper, cited to confirm the work passed peer review rather than remaining a preprint.
- AI Visibility Statistics (2026): How Often Brands Appear in ChatGPT, Perplexity, Claude & Gemini Boring Marketing, 2026 The closest methodological precedent among industry blogs, running repeated brand-visibility audits across four engines; cited for its finding that engines differ substantially in how often they cite a given brand, and flagged as a non-peer-reviewed vendor source.
- ChatGPT search for Enterprise and Edu OpenAI Help Center, 2026 Companion vendor documentation describing the Sources panel, used to describe how citation and ranking are surfaced to end users.
- Gartner: AI Is Reshaping B2B Buying, but Human Sellers Still Close the Confidence Gap Demand Gen Report, 2026 Secondary report of a Gartner buyer survey; cited for the finding that most buyers still validate AI-generated vendor recommendations with a human seller, a caveat on how much a first-place mention actually decides.
- ChatGPT Search OpenAI Help Center, 2026 Vendor documentation confirming ChatGPT search responses can include inline citations, used to describe the mechanism by which a named company enters an answer.
- 100 Most Cited Domains in ChatGPT Ahrefs, 2026 Single-snapshot ranking of ChatGPT's most-cited domains via Ahrefs Brand Radar, used as a point-in-time contrast to our repeated-measures approach.
- AI Platform Citation Patterns: How ChatGPT, Google AI Overviews, and Perplexity Source Information Profound (tryprofound.com), 2026 Cross-engine comparison of citation concentration by source type, used to note that engines differ systematically in whom they cite, a factor relevant to why leader identity might vary by engine.
- 94% of B2B Buyers Use AI for Vendor Research (Forrester 2026 Buyers' Journey Survey, reported) Machine Relations, 2026 Secondary report of a Forrester buyer survey; cited for figures on how commonly AI tools are used to compare vendors before contact, establishing why first-place mentions matter commercially.
- How ChatGPT sources the web Profound (tryprofound.com), 2026 Large-sample study of ChatGPT citation behavior over a defined quarter, cited for its methodology description and for cross-platform citation comparisons.
- GEO: Generative Engine Optimization Princeton University, Collaborate repository, 2024 Institutional repository record confirming authorship and affiliation of the GEO paper.
- Web search tool Anthropic, Claude Platform Docs, 2026 Vendor documentation describing how Claude's web search tool retrieves and cites live web content, used to explain why Claude's answers can change week to week as underlying search results change.
- Half of B2B Software Buyers Now Start Their Research with AI Chatbots G2, via Demand Gen Report, 2026 Secondary report of a G2 buyer survey; cited for the finding that buyers think more highly of vendors an AI chatbot recommends, supporting the commercial stakes of leader-position volatility.
- Citations Anthropic, Claude Platform Docs, 2026 Documentation of how Claude chunks source text for citation, used to explain the mechanics of citation granularity referenced in our discussion of method limits.
Data
The complete row-level dataset is published open and ungated under CC BY 4.0. Every number in this study can be recomputed from it.
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
