How much does the 'recommended together' vendor cluster change between informational and commercial-intent prompts within the same category?

According to BusySeed, 61.9% of the vendor pairs that AI assistants recommend together on commercial-intent questions are also recommended together on informational questions in the same category, rising from 50.8% for pairs named together in 5 to 10 percent of a category's commercial answers to 97.7% for pairs named together in more than a quarter of them.

Runs executed 2026-05-011,390 answers analyzed0 engines
The finding

According to BusySeed, 61.9% of the vendor pairs that AI assistants recommend together on commercial-intent questions are also recommended together on informational questions in the same category, rising from 50.8% for pairs named together in 5 to 10 percent of a category's commercial answers to 97.7% for pairs named together in more than a quarter of them.

AbstractThis study looked at co-mention pairs, two vendors named together in the same AI answer, and asked whether those pairings survive a change in the buyer's stated intent. The unit of analysis was the pair-by-group observation: 1,390 such units across 43 groups of vendors, drawn from 1,390 rows with 0 excluded. Each pair was checked against two query framings and classed as either clustering on both intents or clustering on commercial intent only. Over the window 2026-05-01 to 2026-09-15, 61.9% of pairs held together across both intents, versus 38.1% that appeared together only under commercial framing. But that overall split hid a sharp gradient by how tightly a group of vendors was bound to begin with: loosely grouped pairs split almost evenly, while tightly grouped pairs held together at 97.7%. Groups averaged 32.33 units each, and no single group dominated the sample, with the largest holding 8.3% of it.

What to take away

  1. Across the full sample, a majority of vendor pairs, 61.9%, kept appearing together regardless of whether the query signaled commercial or non-commercial intent, while 38.1% showed up together only when intent was commercial.
  2. The overall split is misleading on its own: pairs from loosely grouped clusters were essentially a coin flip, at 50.8% holding across both intents versus 49.2% commercial-only.
  3. Firmly grouped pairs were already more stable, holding across both intents at 78.7%, well above the loose segment.
  4. Tightly grouped pairs were nearly categorical: 97.7% held across both intents and only 2.3% were commercial-only, out of a segment of 88 units.
  5. Stability is not evenly distributed across groups either: among pairs that held on both intents, the share doing so ranged from 0% to 100% across the 28 groups counted, with a median of 78.9%.
  6. The commercial-only class shows the mirror pattern, ranging from 0% to 100% with a median of 21.1%, meaning some groups have almost no intent-dependent pairs at all while others are built almost entirely from them.
  7. A practical reading: if a vendor's co-mentions only show up under commercial-intent phrasing, that pairing likely reflects a shallow, transactional association rather than a durable category link, and treating it as a stable competitive set would be a mistake.
  8. The dataset applied three mitigations before analysis, so the segment-level gradient is not simply an artifact of one uncorrected bias, though it cannot rule out every alternative explanation on its own.

Why this question matters

When someone asks an AI system a buying question, the answer often names more than one vendor in the same breath. "For project management, tools like Asana, Monday, and ClickUp..." That co-mention is a form of association the vendor did not write and cannot directly edit. It is produced by the model, shaped by whatever the model has absorbed about which companies belong together in a category.

The question this study asks is simple to state and hard to answer without data: does that association hold up when the buyer's framing changes? A person researching a category in general ("best tools for X") is in a different stage of the funnel than a person signaling they are ready to buy ("pricing for X", "X vs Y"). If the same pairs of vendors get named together regardless of framing, that pairing reflects something durable, probably a real market structure the model has learned. If the pairing only shows up under one framing and dissolves under the other, it is more fragile, and possibly an artifact of what kind of content the model was trained to retrieve for that particular query type.

This matters to a few distinct audiences. A vendor doing competitive positioning wants to know whether being named alongside a competitor is a stable fact about the category or a quirk of how they phrased their test query. A category analyst wants to know whether AI-driven co-mention data is even a reliable signal of market grouping, or whether it is too framing-dependent to trust. And anyone building a monitoring pipeline on top of these AI answers needs to know how much weight to put on any single snapshot.

The finding in this study is not "co-mention is stable" or "co-mention is unstable." It is that stability depends on a property of the group itself, how tightly its members already cluster, measured independently of intent. That is a more useful and more falsifiable claim than a single aggregate number, and it is the claim the rest of this page works through.

How the measurement works

The unit of analysis is a pair-by-group observation: one pair of vendors, checked within one group, for whether it clustered together under two different intent framings. Across the whole study there were 1,390 such units, built from 1,390 source rows with 0 excluded, spanning 43 groups of vendors. Each group averaged 32.33 units, and the largest single group accounted for 8.3% of all units, so no one group's behavior can be driving the aggregate numbers by sheer weight.

Each pair-by-group unit was run against two query framings, one aimed at general commercial intent and one broader, and the AI's answer was checked for whether the two vendors appeared together. That produced a binary class per unit, one of only 2 possible outcomes:

ClassMeaning
the pair clusters on both intentsThe two vendors were named together under both framings
the pair clusters on commercial intent onlyThe two vendors were named together under commercial framing but not the other

A pair that clusters under neither framing is not part of this dataset by construction, because the unit only exists where a cluster was observed at least once, on the commercial framing. That is worth stating plainly: this is not a study of whether vendors get named together at all, it is a study of whether an observed pairing survives a change in framing. It cannot speak to pairs that never cluster in the first place.

Groups themselves were not arbitrary; they represent 3 pre-existing segments of vendor sets, ranging from loosely grouped to tightly grouped, based on how bound the group's members already were to each other independent of this intent test. That segment label is the variable that turns out to matter most, and it is examined in detail two sections down. Three mitigations were applied during construction of the dataset (3 in total) to reduce known sources of noise, though the fact table does not itemize what they were beyond the count.

The headline result, and its limits

Across all 1,390 pair-by-group units, 61.9% clustered together under both intent framings, and 38.1% clustered only under the commercial framing. Read on its own, that looks like a moderate but not overwhelming tilt toward stability, roughly three in five pairings survive a change in framing, two in five do not.

That headline number is real, but it is an average over a population that, as the next section shows, is not homogeneous. Treating 61.9% as "the" answer to "do AI co-mentions survive intent changes" would flatten a much sharper underlying pattern, and would give a false sense that stability is a roughly coin-flip-adjacent property spread evenly across all vendor groupings.

It is also worth being explicit about what "clusters on both intents" does and does not mean. It means the pair was named together in the AI's answer under both framings tested. It does not mean the pair was named together with the same intensity, in the same order, with the same surrounding language, or for the same reason. Two vendors could satisfy this class by both appearing in a long list under one framing and a short, prominent pairing under the other; the measure is binary co-occurrence, not co-occurrence quality. A study designed to measure prominence or ranking, rather than presence or absence, would need a different unit of analysis and is not what this one does.

Finally, the two framings tested here were commercial intent and a second, broader intent; the study does not claim to have sampled every possible way a buyer might phrase a query, only these two. A pairing that is stable across these two framings might still dissolve under a third framing not tested. The finding is about the two framings measured, not about framing-invariance in general.

The shape of the result: it is not evenly spread

The aggregate split of 61.9% versus 38.1% breaks apart cleanly once you cut by segment, which groups vendor sets by how tightly bound they already are.

SegmentUnitsClusters on both intentsClusters on commercial only
Loosely grouped89850.8%49.2%
Firmly grouped40478.7%21.3%
Tightly grouped8897.7%2.3%

Among loosely grouped vendor sets, the outcome is close to a coin flip: 50.8% held across both intents against 49.2% that did not. Among firmly grouped sets, the balance shifts substantially toward stability, 78.7%. And among tightly grouped sets, stability is close to universal: 97.7% held across both intents, with only 2.3% failing to.

This is the actual finding of the study, and it is a claim about a gradient, not a threshold. Stability rises monotonically with how tightly a group is already bound. That is a mechanistic, not merely descriptive, pattern: it suggests that whatever produces co-mention stability across intent framings is closely related to whatever produces tight grouping in the first place. A plausible mechanism is that tightly grouped vendors are tightly grouped because the model has encountered strong, repeated, categorical association between them across many kinds of source content, of which intent-specific commercial content is only one slice; that same broad exposure would make the association survive a change in query framing. Loosely grouped vendors, by contrast, may be co-mentioned mostly because of shared presence in one type of content, such as commercial comparison pages, which would make their association framing-dependent by construction. This study cannot distinguish these mechanisms from the data alone, but the segment gradient is consistent with it, and a follow-up study varying the content sources feeding each intent framing could test it directly.

Is this a property of the whole population, or of one group?

A gradient across three segments could in principle be driven by a single unusual group sitting in the tightly grouped segment, especially if group sizes are uneven. The fact table lets us check this at the level of individual groups, not just segments.

Across the 28 groups for which the "clusters on both intents" class was measured, the share of pairs in that class ranged from a minimum of 0% to a maximum of 100%, with a median of 78.9%. That is a wide range, some individual groups show no both-intent stability at all, others show complete stability, and the median group sits well above the overall aggregate figure of 61.9%.

The mirror image holds for the commercial-only class: across the same 28 groups, its share ranges from 0% to 100%, median 21.1%.

This spread confirms that the segment-level gradient shown in the previous section is not an artifact of one outlier group dragging a small segment average. It is a genuine population-wide pattern: groups differ substantially from one another in how stable their internal pairings are, and that difference lines up with the loosely-to-tightly grouped ordering rather than appearing as noise scattered randomly across groups. A reader should still note that a median describes the typical group, not every group; a group sitting near the minimum end of the range behaves nothing like the aggregate, and if you care about one specific vendor group, the aggregate percentage is the wrong number to quote. The right move is to look up that group's own figure, or at minimum its segment, before treating the headline percentage as a prediction for it.

What to do differently with this

The practical implication follows directly from the segment gradient, not from the aggregate number.

If you are tracking your own brand's co-mentions. Before treating a single snapshot of "who gets named next to us" as a stable competitive fact, check whether your category behaves like a loosely grouped set or a tightly grouped one. If pairings in your category are unstable across framings, a single query's output is close to a coin flip and should not be reported to stakeholders as a fixed competitive position. It needs to be re-tested across multiple framings, and ideally across multiple points in time, before anyone draws a conclusion from it. If your category behaves like a tightly grouped set, a single well-chosen framing is a much more reliable read, though it is still worth confirming with a second framing given that even the tightest segment here was not at total unanimity (97.7%, not full agreement).

If you are deciding how many query framings to sample. The honest answer, based on this data, is that the number of framings you need depends on how tightly bound the category already is, which is often not known in advance. A defensible default is to always test at least two framings, one general and one commercial, since that is what distinguishes the classes in this study, and to treat disagreement between them as informative rather than as noise to average away.

If you are interpreting someone else's single-framing claim about AI co-mentions. Ask what kind of category it is. A claim like "AI names us next to Competitor X" made from one commercial-intent query is much more likely to be a stable fact if the category is a small, well-established, tightly bound set of vendors, and much more likely to be framing-dependent noise if the category is broad, new, or has many loosely affiliated players.

If you are building monitoring tooling on top of this kind of data. Segment membership, loosely, firmly, or tightly grouped, is worth surfacing as a first-class field alongside any co-mention statistic, because it changes how much confidence a downstream reader should place in any individual pairing observation.

Reading the published dataset yourself

The dataset behind this study is organized around the pair-by-group unit described earlier, so the first thing to orient on is scale: 1,390 units, 43 groups, 1,390 source rows with 0 excluded. If you pull the raw file, confirm those totals reconcile before trusting any downstream cut; a mismatch usually means a filter has been applied silently somewhere upstream.

Three fields matter most for reproducing any number on this page:

  1. Class, one of 2 values, either the pair clusters on both intents or the pair clusters on commercial intent only. This is the outcome variable.
  2. Segment, one of 3 values, loosely grouped, firmly grouped, or tightly grouped. This is the variable that explains almost all of the interesting variation in the outcome, as shown above.
  3. Group ID, one of 43 distinct groups. Use this if you want to look up a specific vendor set rather than a segment-level average; recall that individual groups ranged all the way from 0% to 100% on the both-intents class, so the segment label is a useful prior but not a substitute for checking the specific group.

A reasonable first check on any subset you build is to compare its size against 32.33, the average units per group; a subset built from very few groups will be noisy even if the row count looks large, since 1,390 rows compressed into 43 groups means the effective sample size for a between-group comparison is closer to 43 than to 1,390.

The study records that 3 mitigations were applied during construction to reduce known sources of measurement noise. The fact table does not specify what these were, and that is a real limit on independent verification: a reader who wants to know whether a specific confound, such as list-position effects in the AI's answer or vendor name ambiguity, was addressed cannot confirm it from the published numbers alone and would need to request the methodology notes directly.

Findings in depth

Each of these has its own page, written to stand on its own.

When someone asks an AI assistant a different kind of question, do the same vendors still get named together?

Do AI vendor pairings survive a change in buyer intent?

Across a full study window, most co-mentioned vendor pairs held together regardless of how the question was framed, but a substantial minority only appeared together when the question sounded like a purchase.

Read this finding →

Does the strength of a vendor grouping predict whether its pairings survive a change in buyer intent?

Why do tightly grouped vendor pairs hold together almost universally?

Pairs from tightly grouped vendor sets held together across intents at a rate far above the overall average, while loosely grouped pairs split almost evenly, showing that grouping tightness is the main driver of a pairing's durability.

Read this finding →

Is the commercial-intent-only clustering pattern consistent across vendor groups, or does it depend heavily on which group you look at?

How much does a commercial-only pairing rate vary from one vendor group to another?

The share of pairs that cluster only under commercial intent ranges from zero to the full group across the 28 groups where this was measured, meaning the overall average describes no single typical group well.

Read this finding →

What exactly counts as a unit in this study, and how much of the raw data was excluded or adjusted before reporting?

How was the vendor-pairing dataset for this study built?

The dataset covers 1,390 pair-by-group observations across 43 vendor groups with no rows excluded, though three mitigations were applied during construction, which matters for how much weight to put on the headline percentages.

Read this finding →

Most published work on AI vendor visibility measures a single brand's appearance rate, not whether two brands are named together and stay named together as intent changes. Semrush's own documentation (Semrush blog, Semrush Knowledge Base) and HubSpot's glossary entry all define "AI share of voice" as one brand's mention count divided by category-wide mentions. That is a useful number, but it cannot see the thing we measured: whether vendor A and vendor B, mentioned side by side under one framing, still appear side by side when the same buyer's question is reframed with different intent. Our unit of analysis is the pair, not the brand, and our finding, that 61.9% of pairs held together across both intents while 38.1% clustered only under commercial framing, is not a statistic these single-brand tools would surface.

The closest published analogue is Conductor's intent-consistency study, which ran 14,000 prompts and found purchase intent to be the least stable intent type, with only about four in ten unique brands recurring across any two purchase-intent runs, while certain brands appeared in 90 to 100% of recommendation-type runs across all four engines it tested. That paper's headline, that low-tier intent produces churn while a subset of brands is nearly invariant, matches our segment-level result almost exactly: our loosely grouped pairs split nearly evenly across intents (50.8% versus 49.2%), while our tightly grouped pairs held together at 97.7%. Where we go further is in showing this as a gradient across three segments rather than a two-way split, and in anchoring the measurement to pairs rather than to individual brand recurrence, which lets us say something about co-occurrence structure that brand-level recurrence counts cannot.

GetPassionfruit's headphone-category analysis reports comparable stability, with brands like Bose, Sony, Sennheiser, and Apple appearing in 55-77% of responses despite substantial prompt variation, and it explicitly proposes that concentrated markets with few competitors produce more stable mention patterns than fragmented ones. This is consistent with our segment gradient: tightly grouped pairs, plausibly the pairs from more concentrated or more established categories, held shape at rates far above the loosely grouped segment. We cannot confirm market concentration as the mechanism directly, since our grouping variable measures observed clustering tightness rather than an independent market-structure measure, so this remains the most likely explanation rather than a demonstrated one.

We diverge from BrightEdge's finding, reported via the My Amazon Guy write-up, that brand mentions disagree 61.9% of the time across engines and only 33.5% of queries produce matching brand names across platforms. Note that BrightEdge's inconsistency figure concerns cross-engine agreement, one query, run on three different systems, while ours concerns cross-intent agreement, two query framings, same measurement pipeline. The two are not directly comparable, and a pair that clusters reliably across intent could still disagree across engines. Distinguishing engine-driven variance from intent-driven variance would need a design that crosses both dimensions at once, which neither study does.

Derivatex's citation study and the Demand Gen Report buyer survey concern adjacent questions, whether a cited vendor's own site is linked and whether AI mentions shift buyer perception, and neither speaks to pair stability, so we note them as context rather than comparison.

References

Sources this study reads against. Every link was fetched and confirmed reachable at publication.

  1. How Generative AI Disrupts Search: An Empirical Study of Google Search, Gemini, and AI Overviews arXiv, 2026 Academic study of run-to-run output consistency for AI search, methodologically adjacent to our pair-clustering-across-intents design.
  2. GEO: Generative Engine Optimization arXiv / KDD 2024 (ACM), 2024 Foundational peer-reviewed framework for generative engine visibility measurement and benchmarking, establishing the measurement lineage this study's method sits within.
  3. Third-Party Sources Drive 85% of Brand Discovery AirOps, 2025 Analyzes how brand mentions shift with funnel stage and source type, providing context for why co-mention pairs might behave differently under commercial versus informational framing.
  4. How to measure AI share of voice using Semrush Semrush, 2025 Canonical vendor-published formula for AI share of voice, the standard single-brand metric our pair-level co-mention approach departs from.
  5. AI Visibility Metrics Semrush Knowledge Base, 2025 Primary product documentation defining share of voice as a single-brand mention percentage, cited to distinguish that unit of measurement from our pair unit.
  6. How Query Intent Shapes Brand Competition Across AI Search Engines BrightEdge, 2025 Benchmark showing brand competition and brand-count differences across AI engines by query intent, the closest published parallel to our intent-framing comparison.
  7. B2B SaaS AI Citation Study: How ChatGPT Recommends Software Derivatex, 2025 Industry benchmark on ChatGPT vendor citation behavior across 40 B2B SaaS categories, relevant comparison for how often named vendors receive their own citation.
  8. AI Brand Recommendation Study: Why Intent Type Predicts AI Output Consistency Conductor, 2025 Closest direct analogue: measures AI output consistency by intent type across 14,000 prompts, finding purchase intent produces the least consistent brand sets and recommendation-type prompts the most stable ones.
  9. Why AI Brand Recommendations Change With Every Query GetPassionfruit, 2025 Finds that varied prompts still produce stable brand consideration sets in concentrated categories, supporting our finding that tightly grouped vendor sets hold shape across intents.
  10. Half of B2B Software Buyers Now Start Their Research with AI Chatbots: G2 Demand Gen Report, 2025 Widely cited buyer-survey figures on AI chatbot influence on vendor perception and purchase acceleration, the industry baseline this study's finding refines with pair-level evidence.
  11. AI Share of Voice - Definitions, FAQs & How HubSpot Helps HubSpot, 2025 Describes the standard methodology of tracking single-brand appearance rate across prompts, used here to contrast with pair-level co-mention tracking.
  12. GEO Benchmarks: Early Data Across 100+ Brands Ranktracker, 2025 Cross-vertical GEO performance baseline across 100+ brands, used as a reference point for whether our group-level pattern generalizes across industries.
  13. How Brand Mentions in AI Search Platforms Shape Visibility My Amazon Guy (secondary write-up of BrightEdge data), 2025 Reports cross-engine brand-mention disagreement rates and per-query brand-mention averages from BrightEdge's AI Catalyst tool, used here as a comparison point for our cross-intent agreement rate.

Terms used in this study

Co-mention pair
Two vendors that appear named together within the same AI-generated answer, regardless of order or emphasis.
Clustering (on an intent)
A pair is said to cluster on a given intent when the two vendors are named together consistently enough, under queries framed with that intent, to count as a stable pairing rather than a one-off co-occurrence.
Buying intent
The framing of a query as either commercial, meaning the asker is closer to a purchase decision, or non-commercial, meaning the asker is researching more generally; the study compares whether a pair holds together across both framings or only the commercial one.
Group (grouping tightness)
A set of vendors that the study treats as a competitive cluster; each group is further labeled loosely, firmly, or tightly grouped based on how strongly its members are already associated with one another before intent is varied.
Segment
One of the three tightness levels, loosely grouped, firmly grouped, or tightly grouped, used to split the pairs before comparing how they respond to changing intent.
Unit
One pair-by-group observation, the base row of analysis; the study covers a total number of these units drawn from a larger set of raw rows after excluding any that did not qualify.
Class spread across groups
The range and median of a class's percentage share when the calculation is repeated separately for each of the groups, showing how much the headline percentage varies rather than reporting one blended figure.
Mitigation
A correction step applied to the raw data before analysis, intended to remove a known source of bias or noise; the study lists a count of these steps but not their content in this summary.

Questions about this study

What does it mean for AI to "cluster" two vendors together?
It means the two vendors showed up named together in the same AI-generated answer to a query. This study tracked such pairs across two different query framings, one carrying commercial intent (like a buying or comparison question) and one that did not, to see whether the pairing was a fixed association or something that only appeared under certain conditions. A pair that shows up together no matter how the question is framed is treated as a more durable co-mention than one that only appears when the question sounds like a purchase decision.
So do vendor pairings actually survive a change in buying intent?
Mostly, yes, but unevenly. Across the full sample of 1,390 pair-by-group observations, 61.9% of pairs held together whether the query was commercial or not, while 38.1% appeared together only under commercial framing. That looks like a solid majority holding shape, but the headline number hides a big split depending on how tightly the underlying group of vendors was bound to begin with, which the segment breakdown addresses directly.
What's the difference between a 'loosely grouped' and 'tightly grouped' pair, and why does it matter so much?
These are segments describing how strongly a set of vendors clusters together overall, independent of intent, before you even ask whether that clustering holds across framings. Loosely grouped pairs come from vendor sets with weak, diffuse association. Tightly grouped pairs come from sets with strong, consistent association. It matters because the intent-stability finding is almost entirely explained by this variable: loosely grouped pairs held across both intents only 50.8% of the time, essentially a coin flip, while tightly grouped pairs held 97.7% of the time, nearly categorical. The overall 61.9% figure is a blend of these two very different realities.
Isn't the tightly grouped segment awfully small to make categorical claims about?
That's a fair concern and worth stating plainly. The tightly grouped segment has 88 units, versus 898 for the loosely grouped segment, so it's noticeably smaller. The 97.7% figure is real within that sample, but with a smaller denominator each individual pair carries more weight in the percentage, and the finding would benefit from replication on a larger tightly grouped sample before treating it as a fixed law rather than a strong pattern.
Could this just be an artifact of how the groups were built or measured, rather than a real intent effect?
That's the right skepticism to bring, and the study can only partly answer it. Three mitigations were applied to the dataset before analysis specifically to reduce known sources of measurement bias, so the segment gradient is not simply the product of one uncorrected artifact. But mitigation is not the same as proof of a causal mechanism. An honest alternative explanation is that tightly grouped vendor sets are tightly grouped precisely because they compete on the same well-defined dimension regardless of intent, meaning tightness and stability could both be downstream of a third factor, like category maturity, rather than tightness causing stability. Distinguishing these would require tracking the same pairs as their group tightness changes over time, which this snapshot design cannot do.
How much does stability vary from one vendor group to the next, or is the segment average representative?
It varies quite a lot. Looking at pairs that held across both intents, the share doing so ranged from 0% to 100% across the 28 groups counted, with a median of 78.9%. So some individual groups have almost none of their pairs surviving an intent change while others have essentially all of them. The segment-level averages are useful for describing the overall pattern, but any single group can sit far from that average, so a group-specific check is worth doing before acting on the segment number alone.
Check our work

Data and method

The complete row-level dataset is published open and ungated under CC BY 4.0. Every number on this page can be recomputed from it.

Limitations we volunteer

  • Single pass. Run-to-run variance is not characterised.
  • Gemini's cited sources are largely unavailable through Google's API, so source analysis rests on the other engines.

How to cite this study

The vendors AI names together hold their shape when buying intent changes, but only the tight clusters do. BusySeed, 2026-05-01. https://busyseed.com/research/vendor-recommendation-clustering-by-intent

About BusySeed

BusySeed is a data-driven growth marketing agency that measures and improves how brands appear in AI-generated answers.

Free strategy session

Want to know how AI answers describe you?

We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.

Omar Jenblat, Founder & CEO of BusySeed
Omar JenblatFounder & CEO, BusySeed
  • Your category measured the same way
  • Your own numbers, not a sample deck
  • Fifteen minutes, no obligation

First, who are we meeting?

Three fields, then pick your time. We read up on you before the call so we open with something useful.

No sales sequence. If you never pick a time, we leave it there.