Questions run
Purchase-intent questions written the way buyers phrase them, not the way keyword tools suggest.
Which brands get recommended, where the answers get their evidence, and how much the six engines disagree with each other.
Published 4 August 2026 · Data collected May — July 2026 · 18 min read
Purchase-intent questions written the way buyers phrase them, not the way keyword tools suggest.
ChatGPT, Claude, Gemini, Perplexity, and AI Overviews, on a fixed daily schedule.
From analytics and CRM to logistics and compliance, chosen for varied buying-committee sizes.
Answers are non-deterministic, so every figure is a mean across repeated daily captures.
Fewer than one in five citations pointed at a page owned by a vendor in the category. Review sites and forums carried the rest.
Only a third of questions produced the same top recommendation across all six. Tracking one engine tells you little about the others.
Top-three Google rankings made a brand only 2.4× likelier to be named — a weaker link than most teams assume.
Between May and July 2026 we put 12,400 purchase-intent questions to six engines. The questions came from 41 B2B categories, written the way a buyer types rather than the way a keyword tool suggests. Every answer was captured in full and classified by whether it named a brand, cited a source, or made an outright recommendation.
What came back doesn’t look much like search.
The most consistent finding across all six engines was how little the models lean on vendor-owned pages. A brand’s own website — the thing most marketing teams spend their entire budget on — accounted for under a fifth of all citations.
Share of total citations by source category. Percentages exceed 100 where a single answer cited more than one source type.
Within that 19%, one page type did disproportionately well: comparison pages. Vendor pages that directly compared the brand against named alternatives were cited 3.2 times more often than product or feature pages on the same domain.
The single highest-leverage page a brand can publish isn’t a product page. It’s an honest comparison against the alternatives buyers are already weighing.
We expected some variance. We didn’t expect this much. For each question we recorded which brand each engine named first, then counted how many engines agreed.
Distribution of agreement across the six engines. In 69% of questions at least one engine named a different brand first than the majority.
The practical consequence is uncomfortable for anyone monitoring a single engine. If you’re checking ChatGPT alone, you’re seeing one reading of your category — and in roughly two out of three questions, at least one other engine is telling buyers something different.
The engines also differ in how willing they are to recommend at all. Some readily name a single best option; others hedge with a list.
| Engine | Named a brand | Cited a source | Made a pick |
|---|---|---|---|
| ChatGPT | 94% | 83% | 58% |
| Claude | 91% | 44% | 39% |
| Google gemini | 88% | 72% | 47% |
| Perplexity | 97% | 94% | 51% |
| Deepseek | 85% | 38% | 44% |
| AI overviews | 79% | 88% | 22% |
Made a pick counts answers that recommended one brand above the others rather than presenting a neutral list.
A focused set of high-intent prompts your customers actually put into AI assistants — not keyword lists.
Mentions, citations, and recommendations across engines, with clear gaps versus competitors.
A ranked action list ordered by how many questions each fix could move — so the next step is obvious.
Evidence-backed findings you can take to leadership, clients, or your content team without translation.
Three things follow from the data, and none of them are what a keyword strategy would tell you.
It matters — comparison pages especially — but four-fifths of the evidence shaping your category's answers sits on domains you don't own. Getting accurately represented on the review sites and forums models actually reach for is not a nice-to-have.
The 31% full-agreement figure means single-engine monitoring will systematically mislead you about where you stand.
The 2.4× lift for top-three Google rankings sounds meaningful until you compare it to the 3.2× lift from simply having a comparison page. Position in a results list is not the same thing as being the answer.
Every question, every classification, and the per-category breakdowns that didn’t fit in the write-up.