Methodology
How much does the 'recommended together' vendor cluster change between informational and commercial-intent prompts within the same category?
Engines and exact versions
A study that does not name the model version it ran against is not reproducible, because the answer changes when the model does.
| Engine | Model version |
|---|
Run window, UTC: [object Object].
Method
The unit of analysis is a pair-by-group observation: one pair of vendors, checked within one group, for whether it clustered together under two different intent framings. Across the whole study there were 1,390 such units, built from 1,390 source rows with 0 excluded, spanning 43 groups of vendors. Each group averaged 32.33 units, and the largest single group accounted for 8.3% of all units, so no one group's behavior can be driving the aggregate numbers by sheer weight.
Each pair-by-group unit was run against two query framings, one aimed at general commercial intent and one broader, and the AI's answer was checked for whether the two vendors appeared together. That produced a binary class per unit, one of only 2 possible outcomes:
| Class | Meaning |
|---|---|
| the pair clusters on both intents | The two vendors were named together under both framings |
| the pair clusters on commercial intent only | The two vendors were named together under commercial framing but not the other |
A pair that clusters under neither framing is not part of this dataset by construction, because the unit only exists where a cluster was observed at least once, on the commercial framing. That is worth stating plainly: this is not a study of whether vendors get named together at all, it is a study of whether an observed pairing survives a change in framing. It cannot speak to pairs that never cluster in the first place.
Groups themselves were not arbitrary; they represent 3 pre-existing segments of vendor sets, ranging from loosely grouped to tightly grouped, based on how bound the group's members already were to each other independent of this intent test. That segment label is the variable that turns out to matter most, and it is examined in detail two sections down. Three mitigations were applied during construction of the dataset (3 in total) to reduce known sources of noise, though the fact table does not itemize what they were beyond the count.
Limitations we volunteer
Written by us, before anyone else found them.
- Single pass. Run-to-run variance is not characterised.
- Gemini's cited sources are largely unavailable through Google's API, so source analysis rests on the other engines.
Terms used in this study
- Co-mention pair
- Two vendors that appear named together within the same AI-generated answer, regardless of order or emphasis.
- Clustering (on an intent)
- A pair is said to cluster on a given intent when the two vendors are named together consistently enough, under queries framed with that intent, to count as a stable pairing rather than a one-off co-occurrence.
- Buying intent
- The framing of a query as either commercial, meaning the asker is closer to a purchase decision, or non-commercial, meaning the asker is researching more generally; the study compares whether a pair holds together across both framings or only the commercial one.
- Group (grouping tightness)
- A set of vendors that the study treats as a competitive cluster; each group is further labeled loosely, firmly, or tightly grouped based on how strongly its members are already associated with one another before intent is varied.
- Segment
- One of the three tightness levels, loosely grouped, firmly grouped, or tightly grouped, used to split the pairs before comparing how they respond to changing intent.
- Unit
- One pair-by-group observation, the base row of analysis; the study covers a total number of these units drawn from a larger set of raw rows after excluding any that did not qualify.
- Class spread across groups
- The range and median of a class's percentage share when the calculation is repeated separately for each of the groups, showing how much the headline percentage varies rather than reporting one blended figure.
- Mitigation
- A correction step applied to the raw data before analysis, intended to remove a known source of bias or noise; the study lists a count of these steps but not their content in this summary.
References
Sources this study reads against. Every link was fetched and confirmed reachable at publication.
- How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews arXiv, 2026 Academic study of run-to-run output consistency for AI search, methodologically adjacent to our pair-clustering-across-intents design.
- GEO: Generative Engine Optimization arXiv / KDD 2024 (ACM), 2024 Foundational peer-reviewed framework for generative engine visibility measurement and benchmarking, establishing the measurement lineage this study's method sits within.
- Third-Party Sources Drive 85% of Brand Discovery AirOps, 2025 Analyzes how brand mentions shift with funnel stage and source type, providing context for why co-mention pairs might behave differently under commercial versus informational framing.
- How to measure AI share of voice using Semrush Semrush, 2025 Canonical vendor-published formula for AI share of voice, the standard single-brand metric our pair-level co-mention approach departs from.
- AI Visibility Metrics Semrush Knowledge Base, 2025 Primary product documentation defining share of voice as a single-brand mention percentage, cited to distinguish that unit of measurement from our pair unit.
- How Query Intent Shapes Brand Competition Across AI Search Engines BrightEdge, 2025 Benchmark showing brand competition and brand-count differences across AI engines by query intent, the closest published parallel to our intent-framing comparison.
- B2B SaaS AI Citation Study: How ChatGPT Recommends Software Derivatex, 2025 Industry benchmark on ChatGPT vendor citation behavior across 40 B2B SaaS categories, relevant comparison for how often named vendors receive their own citation.
- AI Brand Recommendation Study: Why Intent Type Predicts AI Output Consistency Conductor, 2025 Closest direct analogue: measures AI output consistency by intent type across 14,000 prompts, finding purchase intent produces the least consistent brand sets and recommendation-type prompts the most stable ones.
- Why AI Brand Recommendations Change With Every Query GetPassionfruit, 2025 Finds that varied prompts still produce stable brand consideration sets in concentrated categories, supporting our finding that tightly grouped vendor sets hold shape across intents.
- Half of B2B Software Buyers Now Start Their Research with AI Chatbots: G2 Demand Gen Report, 2025 Widely cited buyer-survey figures on AI chatbot influence on vendor perception and purchase acceleration, the industry baseline this study's finding refines with pair-level evidence.
- AI Share of Voice - Definitions, FAQs & How HubSpot Helps HubSpot, 2025 Describes the standard methodology of tracking single-brand appearance rate across prompts, used here to contrast with pair-level co-mention tracking.
- GEO Benchmarks: Early Data Across 100+ Brands Ranktracker, 2025 Cross-vertical GEO performance baseline across 100+ brands, used as a reference point for whether our group-level pattern generalizes across industries.
- How Brand Mentions in AI Search Platforms Shape Visibility My Amazon Guy (secondary write-up of BrightEdge data), 2025 Reports cross-engine brand-mention disagreement rates and per-query brand-mention averages from BrightEdge's AI Catalyst tool, used here as a comparison point for our cross-intent agreement rate.
Data
The complete row-level dataset is published open and ungated under CC BY 4.0. Every number in this study can be recomputed from it.
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
- Your category measured the same way
- Your own numbers, not a sample deck
- Fifteen minutes, no obligation
