Questions about this study
The State of AI Search, July 2026
12 questions answeredMethod challenges included
What is this study actually measuring?
It measures what three AI assistants, ChatGPT, Claude, and Gemini, say when asked buying questions across a wide span of business categories. Over the run window from 1 July 2026 to 31 July 2026, the study issued 22,875 distinct prompts and collected 1,548,342 usable answers, all logged as ok. Each answer was checked against a cohort of 153,251 companies to see which ones got named, in what position, and by which engine. It is not a survey of opinion or brand awareness. It is a direct log of what these three specific engine versions returned during this specific window, so it describes July 2026 behavior, not a permanent property of any company.
So is my company invisible to AI if one engine doesn't mention it?
No, and this is the central finding. 0 of 153,251 companies never appeared on any engine at all, meaning total invisibility was essentially not observed in this cohort. What happens instead is that 96.1% of the cohort was named by at least one engine somewhere, but no single engine names everyone. A company missing from one engine's answer is far more likely to be a coverage gap on that engine than proof the company doesn't exist to AI search. Check the other two engines before concluding you have a visibility problem.
Why don't the three engines agree with each other more often?
They largely don't. Across 515,720 company mentions that at least one engine made, all three engines agreed on the same company only 3.3% of the time, and 88.1% of named instances came from exactly one engine acting alone. This likely reflects differences in each engine's underlying retrieval and training data rather than differences in the companies themselves. The practical implication is that a visibility score built from a single engine is measuring that engine's idiosyncratic opinion, not a market consensus.
Which engine mentions the most companies?
ChatGPT named the highest share of the cohort at 42.8%, followed by Claude at 40.5%, with Gemini lower at 33%. This gap of roughly nine points between ChatGPT and Gemini is large enough that a company's apparent standing can shift noticeably depending on which engine a prospective buyer happens to use. The study can't say why Gemini names fewer companies, whether from a narrower retrieval set, different prompt handling, or a more conservative citation style, only that the gap is consistent across the run.
What does 'average mention position' mean, and why does it matter?
When an engine names several companies in one answer, mention position is the rank order in which a given company appears, first, second, third, and so on. Across 594,486 mentions logged in this run, the average position was 5.15. Being named at all and being named early are different outcomes: a company that appears in almost every answer but always near the bottom of the list is in a materially weaker position than one named less often but consistently first, since buyers reading a list tend to weight earlier mentions more heavily.
Does the average position differ by engine?
Yes, modestly. Claude's average mention position was 5.34, ChatGPT's was 5.23, and Gemini's was 4.83, against an overall average of 5.15 across 594,486 mentions. The spread between engines is small relative to the scale, so treat this as a mild tendency rather than a strong pattern. It would take a much larger gap, or a consistent widening over several monthly runs, before concluding one engine systematically buries or promotes mentions relative to the others.
Does the type of question a buyer asks change whether a company gets mentioned?
Barely. Appearance rate stayed within a narrow band across prompt types: commercial-investigation prompts, the kind someone asks while comparing options before buying, had an appearance rate of 38.9%, informational prompts came in at 38.7%, navigational prompts at 37.1%, and transactional prompts at 38.4%. The practical takeaway is that optimizing content narrowly for one prompt type, say commercial-investigation phrasing, is unlikely to move appearance rate much, since engines seem to draw from a similar underlying pool of companies regardless of how the question is framed.
Isn't 100% visibility in every category a sign the categories are too broad or the method is off?
It's a fair concern worth naming directly. Every category in this run shows a visible_pct of 100, including categories with very few companies, such as Field Service Management Software with only a handful of firms, alongside categories with thousands, such as Digital Marketing. When a small category shows full visibility, that could mean genuine full coverage, or it could mean the category is too narrow to produce a miss in one month of prompting. Distinguishing those two explanations requires watching whether small categories stay at full visibility over several months or whether a miss eventually shows up once prompt variety increases.
What should a company actually do with this information?
Stop treating a single engine's answer as the full picture. Check standing on all three engines, since 88.1% of named instances come from just one engine, meaning a Gemini-only check would miss most of what ChatGPT and Claude are saying. If your company is missing from an answer, check position as well as presence, since the overall average position across 594,486 mentions was 5.15 and appearing late in a list is a weaker outcome than not appearing at all in that one answer. Track this monthly rather than reacting to one run, since engine behavior can shift between windows.
What counts as a 'usable answer' in this study, and what got excluded?
A usable answer is one where the engine returned a response that was logged successfully and could be parsed for company mentions and position, recorded here as a status of ok. This run collected 1,548,342 such answers out of runs across 3 engines using one model version each, all under the label rankxa-warehouse. The fact table shows a status_counts.ok figure equal to the total row count, meaning no answers were dropped for errors or timeouts in this particular run. That does not mean every prompt produced a mention, only that every attempt returned a parseable response.
Could this whole result just be an artifact of which companies happened to be in the cohort?
That's a real limitation worth stating plainly. The cohort here is 153,251 companies, and every finding, the 96.1% visibility figure, the 3.3% agreement rate, all of it, is conditional on which companies were selected and how categories were defined. If the cohort skewed toward companies with a strong existing web presence, visibility rates would run higher than they would for a colder, less-indexed set of companies. The way to test this is to compare results across cohorts built with different selection criteria and see whether the same low agreement pattern holds.
What does 'sole mention' mean, and why is 88.1% of pairs such a big deal?
A named pair here is one company matched to one prompt where at least one engine mentioned it. A sole mention means only one of the three engines made that mention, the other two stayed silent on that same company for that same prompt. 88.1% of all such pairs were sole mentions, meaning fewer than one in ten pairs, specifically 8.6%, had two engines agreeing and only 3.3% had all three agreeing. This matters because it means most individual AI mentions are not corroborated by the other engines, so a single citation should be read as one engine's judgment, not a verified fact about the company.
Free strategy session
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
Omar JenblatFounder & CEO, BusySeed
- Your category measured the same way
- Your own numbers, not a sample deck
- Fifteen minutes, no obligation
