The State of AI Search, July 2026
According to BusySeed, in July 2026, 88.1% of the company recommendations AI assistants made for a buyer question came from only one of three engines, and just 3.3% were made by all three.
According to BusySeed, in July 2026, 88.1% of the company recommendations AI assistants made for a buyer question came from only one of three engines, and just 3.3% were made by all three.
What to take away
- Almost every company in the cohort, 96.1% of 153,251, was named by at least one engine, so the binding constraint on AI visibility is not existence in the index but which engine happens to be asked.
- The three engines rarely agree: only 3.3% of named pairs were mentioned by all three, and 88.1% were named by exactly one engine, which means a single-engine visibility score describes one engine's opinion, not the market's.
- Gemini names companies less often (33%) than ChatGPT (42.8%) or Claude (40.5%), a gap large enough to change a company's apparent standing depending on which engine a buyer happens to use.
- Appearance rate barely moves with what the buyer is trying to do: commercial-investigation prompts (38.9%) and informational prompts (38.7%) land within a couple of points of each other.
- Being named is not the same as being named first: the average mention position across 594,486 mentions is 5.15, and that average differs slightly by engine, from 4.83 on Gemini to 5.34 on Claude.
- Category-level visibility in this run shows no companies shut out entirely, 0 of 153,251 never appeared anywhere, which argues against treating a single missing mention as proof of exclusion.
Key findings
- 88.1% of company recommendations were made by only one of the 3 engines
- 3.3% were made by all 3 engines for the same question
- 5.15 average mention position where they do appear, across 594,486 appearances
- 42.8% of the cohort is visible on ChatGPT
- 40.5% of the cohort is visible on Claude
- 33% of the cohort is visible on Gemini
Why this question matters, and to whom
A buyer typing a question into ChatGPT, Claude, or Gemini is doing something a search engine result page never asked them to do: trusting a single generated answer instead of scanning ten blue links and picking one. If a company's name does not appear in that answer, the company does not get considered, no matter how good its website or its case studies are. This is a different failure mode than a low search ranking, because there is no page two to fall back on and no visible reason the omission happened.
This matters most to companies that sell into considered-purchase categories, where a buyer researches before contacting anyone: software vendors, agencies, consultants, service providers with a sales cycle longer than an impulse buy. It matters less, at least for now, to businesses that win on proximity or price and where the buyer's next step is a map search rather than a chat.
It also matters to anyone who has already been sold an "AI visibility score" or a "generative engine optimization" audit. Those products imply a single stable ranking exists to be optimized. The first finding in this study, that the three engines agree with each other only 3.3% of the time, is a direct test of that premise. A score built from one engine's answers is a measurement of that engine, not of the buyer's experience, because the buyer might just as easily have opened a different assistant and gotten a different list of names.
The practical stakes are ordinary business ones: which vendor gets the first call, which agency gets shortlisted before a request for proposal goes out, which software gets a free trial signup instead of a competitor's. None of this requires AI assistants to replace search entirely. It only requires enough buyers to ask enough questions this way that being systematically absent from the answer becomes a measurable cost, the same way a low organic ranking became a measurable cost for the two decades before this one. This study is an attempt to measure that cost directly, on a monthly cadence, rather than infer it from anecdote.
How the measurement works
The design goal was to reproduce, as closely as an automated study can, what an ordinary buyer sees: a plain question, put to an assistant, with no special prompting to surface a particular vendor. The study ran 22,875 distinct prompts through 3 engines, each represented by one production model configuration, 1 version tracked for ChatGPT, 1 for Claude, 1 for Gemini, over a run window from 2026-07-01 to 2026-07-31 (UTC timestamps). That produced 1,548,342 logged rows, of which 1,548,342 returned a usable answer, a completion rate the study treats as the denominator for every rate reported here.
Prompts were drawn across four intent types, informational (someone learning about a category), navigational (someone looking for a specific known company), commercial-investigation (someone comparing options before a decision), and transactional (someone ready to act). Sorting by intent matters because a company's absence from an informational answer means something different than its absence from a transactional one: the first is a content gap, the second is closer to a lost sale.
Every usable answer was parsed for company mentions and matched against a cohort of 153,251 companies drawn from 153,251 candidate records. A mention was logged with its position in the answer, first, second, and so on, because position is doing real work: a company named sixth in a list of eight is technically visible and commercially almost invisible.
Three corrections, referred to here as mitigations, were applied before publication (3 in total) to remove known sources of noise, such as an engine echoing the prompt's own wording back as a company name. The study does not claim these mitigations catch everything, and a reader auditing the dataset should check the mitigation log against their own category before trusting a null result for any single company.
The headline result, and its limits
Across the full cohort of 153,251 companies, 96.1% were named by at least one engine at least once during the window. Read alone, that number suggests AI visibility is close to universal, and in a narrow sense it is: 0 of 153,251 companies never appeared on any engine at all.
But universal presence somewhere is not the same as usable presence everywhere, and this is the distinction the rest of the study exists to make. A company can be part of the 96.1% while still being invisible on the one engine its actual buyers happen to use, or invisible for the one prompt type, transactional versus informational, that matters to its sales process. The topline figure answers "does this company exist to AI search at all" and answers it almost entirely in the affirmative. It does not answer "does this company show up when it counts," which is a narrower and more useful question addressed in the sections on engine disagreement and prompt type below.
What the headline number cannot see: it cannot see whether a mention was flattering or accurate, only whether the name appeared. It cannot see whether the buyer who read that answer went on to click, call, or buy. And it cannot distinguish a company that was deliberately excluded, for instance because an engine considers it untrustworthy, from one that was simply never in the training or retrieval material an engine drew on for that category. Those two situations call for opposite responses, and the raw appearance rate does not tell a company which one it is in.
The honest reading is that this is a floor, not a ceiling: 96.1% is close to a ceiling effect, meaning most companies clear the low bar of existing to at least one engine, which makes the more discriminating measurements, position, cross-engine agreement, prompt-type sensitivity, the ones worth acting on.
Is visibility gradual or binary
A company deciding what to do next needs to know whether AI visibility behaves like a light switch, on or off, or like a dimmer, a matter of degree. The data says dimmer, and the shape of that gradient matters more than the average.
Start from the fact that 0 of 153,251 companies were entirely absent everywhere. That alone rules out a binary story where most companies are simply excluded. Instead, visibility varies along at least three continuous dimensions: which engine is asked (42.8% on ChatGPT versus 33% on Gemini), where in the answer a mention lands (average position 5.15 across 594,486 mentions, with real spread around that mean), and how many of the three engines agree on a given company (3.3% full agreement, 8.6% two-engine agreement, 88.1% single-engine mentions, out of 515,720 named pairs total).
This has a direct operational consequence. If visibility were binary, on or off, the fix would be binary too: figure out why the engine excluded you and remove the cause. Because it is gradual, the fix is closer to what companies already do for conventional search: incremental improvement in the underlying signal (structured information about what the company does, third-party mentions, consistency of description across the web) that shifts probability of inclusion and position rather than flipping a switch.
It also means a single bad month, or a single missed mention on one engine, is not proof of a systemic problem. A company sitting in the 88.1% of pairs named by exactly one engine could reasonably be named by a second engine next month without anything about the company changing, simply because these are probabilistic systems drawing on retrieval and training material that shifts over time. The gradient shape means the right monitoring cadence is monthly and comparative, not a single snapshot treated as verdict.
What the engines are reading when they decide who to recommend
None of the three engines in this study, ChatGPT, Claude, Gemini, disclose their retrieval process in a way that lets an outside study observe it directly. What can be inferred comes from the pattern of results rather than from documentation.
The first pattern is that the three engines disagree with each other far more than they agree: 3.3% full three-way agreement against 88.1% single-engine mentions, out of 515,720 pairs where at least one engine named a company. If all three engines drew from the same underlying corpus of facts about a company and applied the same judgment, agreement would be much higher. The gap implies each engine is drawing on a materially different slice of information, whether that is training data cutoff, which web sources it weights, or which retrieval index it queries live at answer time.
The second pattern is the difference in raw appearance rate by engine: ChatGPT names a cohort company 42.8% of the time it is asked, Claude 40.5%, Gemini 33%. This is not proof that Gemini is stricter or more conservative in some principled sense; it is equally consistent with Gemini's underlying retrieval surfacing fewer named entities generally, independent of any judgment about quality. Distinguishing those two explanations, a considered filter versus a narrower retrieval surface, would require comparing how each engine handles a category where independent, verified quality differences between companies are already known, which this study does not attempt.
The third pattern is that appearance rate is close to flat across prompt intent, informational at 38.7%, navigational at 37.1%, commercial-investigation at 38.9%, transactional at 38.4%. If engines were reading intent and adjusting which companies to surface accordingly, a bigger spread would be expected. The narrow spread suggests the underlying retrieval step, deciding which companies are candidates for mention at all, happens largely independent of what the buyer is actually trying to do, with intent affecting phrasing more than the candidate set.
Where the engines disagree, and why one score is not enough
The clearest number in this study is also the one most likely to be misused if read in isolation. 88.1% of pairs named by at least one engine were named by exactly one engine and not the other two. Only 3.3% of pairs achieved full three-engine agreement.
| Engines agreeing | Share of named pairs |
|---|---|
| Exactly one | 88.1% |
| Exactly two | 8.6% |
| All three | 3.3% |
This table is the practical argument against buying a single "AI visibility score" from any vendor that measures one engine, or that blends engines into a single number without showing the underlying split. A single-engine score can be accurate about that engine and still describe a small, possibly unrepresentative slice of what buyers actually encounter, because 88.1% of the time the other two engines said something different, whether that was naming a competitor instead or naming no one from the cohort at all.
There are two honest explanations for this disagreement rate, and they call for different responses. One is that the engines are converging on genuinely different judgments about which companies deserve mention, in which case a company should expect its standing to be lumpy by design, strong on one engine, weak on another, indefinitely. The other is that the engines are drawing on different, partially overlapping information about the same companies, in which case the disagreement is closer to noise that would shrink if every engine had access to the same up-to-date facts about a company. This study's design, repeating the same prompts monthly across all three engines, is built to distinguish these two over time: a persistent single-engine advantage across many months looks like judgment, a rotating one looks like noise.
For now, the operating conclusion is narrower: any claim of the form "we rank number one in AI search" should be read as "we rank number one on the one engine that was checked," until the claim specifies which engines were tested and shows the cross-engine spread.
What a company on the wrong side of this should actually do
A company that finds itself in the 88.1% of pairs named by only one engine, or absent on the engine its buyers actually use, has a narrower and more mechanical set of first steps than the phrase "AI optimization" usually implies.
First, identify which engine matters for the actual buyer, not all three equally. A company selling to enterprise IT buyers who default to Copilot-adjacent tools has a different priority than one selling to consumers who default to whatever ships on their phone. This study measures three engines because that is what is measurable, not because all three carry equal weight for every business.
Second, treat position, not just presence, as the metric to move. The average mention position across this study is 5.15 out of 594,486 mentions; a company appearing but consistently late in the list is functionally in the same position as one that does not appear, because most buyers reading a generated answer act on the first few names. Moving from an occasional mention to a consistent early mention is a different, harder problem than moving from zero mentions to one.
Third, treat the absence of a mention as a hypothesis, not a verdict, and check it against the mitigation log and against next month's run before reacting. Because 0 of 153,251 companies were absent on every engine, and because 96.1% were visible somewhere, a single missing month on a single engine is common and often reverses.
Fourth, treat the actual content decisions as ordinary content and public-relations decisions, not a new discipline invented for this moment: consistent, specific, and third-party-verifiable descriptions of what the company does, in the places engines already draw from, are the closest lever available. This study does not have a top-cited-sources table populated for this run, so it cannot yet say which specific sources moved a given engine's answer; that is a gap in this month's data, not a claim that sources do not matter.
How to read the published dataset yourself
The underlying data behind this study is published alongside it, and a reader checking a specific claim, especially about their own company or category, should know what each field means before drawing a conclusion from it.
Start with the run window. Every row carries a UTC timestamp between 2026-07-01T23:15:00Z and 2026-07-31T17:15:00Z, so a company that changed its website, launched a product, or issued a press release mid-month may show different behavior in the first half of the window than the second. The dataset does not currently break out sub-month periods, so this kind of within-month change is invisible unless a reader compares this month's file to next month's.
Each row is one logged prompt-engine pair with a status field; 1,548,342 of 1,548,342 rows are marked ok, meaning a usable answer was returned. Rows that are not ok are excluded from every rate in this study, including the appearance rates and the position averages, so a category or engine with a lower ok rate has a smaller effective sample than its row count suggests, and a reader should check the ok count for any narrow slice before trusting a percentage built on it.
Category-level tables in the dataset (by_category) report a visible count and a visible_pct alongside the raw company count, which lets a reader check whether a category's headline visibility figure is built on a denominator large enough to trust, several categories in this run have companies counts in the single or low double digits, and a percentage from a denominator that small should be treated as suggestive rather than conclusive.
The pair_agreement table is the one most likely to be misread. Its denominator is pairs_named_by_at_least_one_engine, 515,720 in this run, not the full cohort. A company that no engine named at all does not enter this table, so the agreement percentages describe the behavior of engines conditional on at least one of them noticing the company, not the behavior of engines toward the whole cohort. Conflating the two denominators is the single easiest way to misstate this study's findings, and any claim that quotes a pair_agreement percentage should say, in the same sentence, that it is conditional on at least one mention.
Findings in depth
Each of these has its own page, written to stand on its own.
What share of companies actually get mentioned by AI assistants at all?
Almost every company shows up somewhere. That is not the same as showing up where it counts.
Nearly all of the cohort was named by at least one engine during the study window, but that headline figure hides how unevenly that presence is distributed.
Read this finding →If ChatGPT recommends a company, does Claude or Gemini recommend the same one?
The three engines agree with each other on almost nothing.
Out of every pair of engine and company where at least one engine made a mention, the three engines landed on the same answer together only a small fraction of the time.
Read this finding →Does it matter which AI engine a buyer happens to use?
Gemini names companies less often than ChatGPT or Claude do.
Across the same set of prompts, the three engines named a cohort company at noticeably different rates, with Gemini the least likely of the three to name anyone from the cohort.
Read this finding →Once a company is mentioned by an AI assistant, where does it typically land in the list?
Being named is not the same as being named first.
The average position of a mention across all three engines sits in the middle of a typical answer, and the average shifts slightly depending on which engine is asked.
Read this finding →References
Sources this study reads against. Every link was fetched and confirmed reachable at publication.
- GEO: Generative Engine Optimization arXiv (Princeton, Georgia Tech; KDD 2024), 2024 Foundational paper formalizing GEO and introducing GEO-bench, showing content-level interventions can shift visibility in generative engine responses.
- E-GEO: A Testbed for Generative Engine Optimization in E-Commerce arXiv, 2025 Proposes a domain-specific benchmark for studying GEO effects in e-commerce, relevant context for our multi-category commercial cohort.
- Deep-Research Agents Can Be Poisoned via User-Generated Content arXiv, 2026 Literature review section summarizes GEO findings that authoritative language, citations, and statistics affect source selection, and that engines show distinct citation preferences from traditional search.
- Generative engine optimization Wikipedia, 2026 General reference for GEO terminology and definition used to orient readers unfamiliar with the field.
- GEO: Generative Engine Optimization ACM Digital Library, Proceedings of KDD 2024, 2024 Conference of record for the GEO paper; cited for the peer-reviewed venue and abstract.
- ChatGPT Search OpenAI Help Center, 2026 Vendor documentation describing how ChatGPT presents inline citations, used to explain mechanism behind engine-reported sources.
- The Most-Cited Domains in AI: A 3-Month Study Semrush, 2026 Independent large-sample study of citation share across LLMs over thirteen weeks, used as a comparison point for cross-engine citation behavior.
- AI Platform Citation Patterns: How ChatGPT, Google AI Overviews, and Perplexity Source Information Profound (tryprofound.com), 2026 Cross-engine analysis of citation patterns by top-level domain, relevant comparison for engine-level divergence.
- How ChatGPT sources the web Profound (tryprofound.com), 2026 Large-sample (~730,000 conversation) study of ChatGPT citation behavior, used to compare scale and method against our own answer set.
- How B2B Buyers Use AI to Choose Vendors Grey Matter, 2026 Summarizes Forrester's 2026 Buyer Insights survey finding generative AI is now buyers' most-cited research source, used to motivate why engine visibility matters commercially.
- Web search tool Claude Platform Docs, Anthropic, 2026 Vendor documentation of Claude's web search and citation mechanism, cited to explain how Claude selects and attributes sources.
- Anthropic Introduces Web Search Functionality for Claude Models InfoQ, 2025 Independent technical summary of the Claude web search API launch, corroborating the citation-generation mechanism.
- Half of B2B Software Buyers Now Start Their Research with AI Chatbots: G2 Demand Gen Report, 2026 Reports G2's 'Answer Economy' finding on ChatGPT's dominant share of B2B software research starts, used as buyer-behavior context.
- 100 Most Cited Domains in ChatGPT Ahrefs, 2026 Monthly-updated domain-level citation tracker for ChatGPT, cited for its finding on the concentration of citations among a handful of large domains.
- Grounding overview Google Cloud Documentation, Gemini Enterprise Agent Platform, 2026 Defines grounding as tethering model output to verifiable sources, used to explain the general concept applied across all three engines.
- 72% of B2B software buyers now use ChatGPT to evaluate vendors, and most brands aren't showing up MarketScale, 2026 Cites Forrester/Crackle PR Q2 2026 AI Citation Benchmark on buyer use of ChatGPT for vendor evaluation and the share of brands absent from citations, used to frame the stakes of non-visibility.
- Grounding with Google Search Google AI for Developers, 2026 Vendor documentation of Gemini's grounding mechanism, cited to explain how Gemini attaches citations to generated answers.
Terms used in this study
- Appearance rate
- The share of answers, out of all usable answers for a given slice such as an engine or a prompt type, in which at least one company from the cohort was named.
- Cohort
- The fixed set of companies, 153,251 of them in this study, that every engine answer is checked against to determine whether a mention occurred.
- Mention position
- The rank at which a company appears within a single generated answer, first, second, third, and so on; a lower number means the company appeared earlier in the list.
- Pair agreement
- A measure of how many of the three engines named the same company for the same prompt, calculated only over company-prompt pairs where at least one engine made a mention.
- Prompt type
- A classification of the buyer intent behind a prompt: informational (learning about a category), navigational (looking for a known company), commercial-investigation (comparing options), or transactional (ready to act).
- Run window
- The start and end timestamps, in UTC, over which this month's prompts were issued and answers collected; here, 2026-07-01T23:15:00Z to 2026-07-31T17:15:00Z.
- Status: ok
- A row-level flag meaning the engine returned a usable, parseable answer for that prompt; rows not marked ok are excluded from every rate reported in this study.
- Visible somewhere, not everywhere
- The condition of being named by at least one of the three engines during the window, without necessarily being named by all three or named consistently.
- Zero visibility
- The condition of never being named by any engine, on any prompt, during the entire run window; 0 of 153,251 companies were in this state this month.
Questions about this study
What is this study actually measuring?
So is my company invisible to AI if one engine doesn't mention it?
Why don't the three engines agree with each other more often?
Which engine mentions the most companies?
What does 'average mention position' mean, and why does it matter?
Does the average position differ by engine?
Data and method
The complete row-level dataset is published open and ungated under CC BY 4.0. Every number on this page can be recomputed from it.
Limitations we volunteer
- Single pass. Run-to-run variance is not characterised.
- Gemini's cited sources are largely unavailable through Google's API, so source analysis rests on the other engines.
How to cite this study
The State of AI Search, July 2026. BusySeed, 2026-07-01. https://busyseed.com/research/state-of-ai-search-2026-07
About BusySeed
BusySeed is a data-driven growth marketing agency that measures and improves how brands appear in AI-generated answers.
More BusySeed research
The 50 fastest-growing SaaS companies barely exist in AI answers
According to BusySeed, 5 of 50 companies studied (10%) never appear when buyers ask AI assistants about their own category.
Does the Company AI Recommends First Stay the Same? A 15-Week Volatility Study
According to BusySeed, the company an AI assistant names first for a category changes from one week to the next in 56.4% of week-to-week transitions, and 82.3% of the question and engine pairs tracked saw their top answer change at least once.
Do AI Engines Recommend the Same Companies? A 3-Engine Agreement Study
According to BusySeed, 88.4% of the company recommendations AI assistants make for a buyer question come from only one of three engines, and just 3.2% are made by all three.
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
- Your category measured the same way
- Your own numbers, not a sample deck
- Fifteen minutes, no obligation
