AEO report — Q2 2026

We ran 12,400 buyer questions through six AI engines.

Which brands get recommended, where the answers get their evidence, and how much the six engines disagree with each other.

Read the findings

Published 4 August 2026 · Data collected May — July 2026 · 18 min read

At a glance

12,400

Questions run

Purchase-intent questions written the way buyers phrase them, not the way keyword tools suggest.

6

Engines tracked

ChatGPT, Claude, Gemini, Perplexity, and AI Overviews, on a fixed daily schedule.

41

B2B categories

From analytics and CRM to logistics and compliance, chosen for varied buying-committee sizes.

90

Days of runs

Answers are non-deterministic, so every figure is a mean across repeated daily captures.

Key findings

19%

Vendor pages barely get cited

Fewer than one in five citations pointed at a page owned by a vendor in the category. Review sites and forums carried the rest.

31%

The engines rarely agree

Only a third of questions produced the same top recommendation across all six. Tracking one engine tells you little about the others.

2.4×

Google rank doesn't transfer

Top-three Google rankings made a brand only 2.4× likelier to be named — a weaker link than most teams assume.

Every team we spoke to while building Heft asked a version of the same question: is AI search actually different, or is it search with a chat interface on top? So we ran the experiment.

Between May and July 2026 we put 12,400 purchase-intent questions to six engines. The questions came from 41 B2B categories, written the way a buyer types rather than the way a keyword tool suggests. Every answer was captured in full and classified by whether it named a brand, cited a source, or made an outright recommendation.

What came back doesn’t look much like search.

Where answers get their evidence

The most consistent finding across all six engines was how little the models lean on vendor-owned pages. A brand’s own website — the thing most marketing teams spend their entire budget on — accounted for under a fifth of all citations.

Citation sources by typeAll engines · 12,400 questions
  • Review sites
    41%
  • Forums
    27%
  • Vendor sites
    19%
  • News & press
    13%

Share of total citations by source category. Percentages exceed 100 where a single answer cited more than one source type.

Within that 19%, one page type did disproportionately well: comparison pages. Vendor pages that directly compared the brand against named alternatives were cited 3.2 times more often than product or feature pages on the same domain.

The single highest-leverage page a brand can publish isn’t a product page. It’s an honest comparison against the alternatives buyers are already weighing.

How much the engines disagree…

We expected some variance. We didn’t expect this much. For each question we recorded which brand each engine named first, then counted how many engines agreed.

Engines agreeing on the top recommendation

Per question
4%
1 engine
11%
2 engines
19%
3 engines
24%
4 engines
11%
5 engines
31%
All 6

Distribution of agreement across the six engines. In 69% of questions at least one engine named a different brand first than the majority.

The practical consequence is uncomfortable for anyone monitoring a single engine. If you’re checking ChatGPT alone, you’re seeing one reading of your category — and in roughly two out of three questions, at least one other engine is telling buyers something different.

Per-engine behaviour

The engines also differ in how willing they are to recommend at all. Some readily name a single best option; others hedge with a list.

Answer behaviour by engine

Share of answers
EngineNamed a brandCited a sourceMade a pick
ChatGPT94%83%58%
Claude91%44%39%
Google gemini88%72%47%
Perplexity97%94%51%
Deepseek85%38%44%
AI overviews79%88%22%

Made a pick counts answers that recommended one brand above the others rather than presenting a neutral list.

What’s inside the report

01

The questions buyers ask

A focused set of high-intent prompts your customers actually put into AI assistants — not keyword lists.

02

Where you stand today

Mentions, citations, and recommendations across engines, with clear gaps versus competitors.

03

What to fix first

A ranked action list ordered by how many questions each fix could move — so the next step is obvious.

04

Proof you can share

Evidence-backed findings you can take to leadership, clients, or your content team without translation.

What this means if you’re the brand

Three things follow from the data, and none of them are what a keyword strategy would tell you.

Your own site is the smaller lever.

It matters — comparison pages especially — but four-fifths of the evidence shaping your category's answers sits on domains you don't own. Getting accurately represented on the review sites and forums models actually reach for is not a nice-to-have.

One engine is not a proxy for the rest.

The 31% full-agreement figure means single-engine monitoring will systematically mislead you about where you stand.

Search rank is a weak signal.

The 2.4× lift for top-three Google rankings sounds meaningful until you compare it to the 3.2× lift from simply having a comparison page. Position in a results list is not the same thing as being the answer.

How we ran this

  • 12,400 purchase-intent questions across 41 B2B categories, written as buyers phrase them.
  • Each question run on all six engines on a fixed daily schedule, May to July 2026.
  • Answers captured as a user sees them, not via thinner API variants.
  • Every figure is a mean across repeated daily runs, not a single capture.
  • Full methodology, question set and raw counts are in the downloadable dataset.

Get the full dataset

Every question, every classification, and the per-category breakdowns that didn’t fit in the write-up.

Check your own brand
AEO report — Q2 2026 | Heft®