How ChatGPT, Claude, Perplexity, and AI Overviews Retrieve and Cite: The Engine Mechanics Reference
ChatGPT layers citations onto an answer. Perplexity builds the answer from citations. The four engines are not one surface.

The four engines that matter retrieve differently, cite differently, and reward different things. ChatGPT layers citations onto an answer it was going to write anyway. Perplexity builds the answer from citations. Claude sits closer to research-stage questions and sends a surprising share of B2B referral traffic for its user base. Google AI Overviews is gated by ordinary Search indexing and mostly does not send anyone anywhere. Optimising for all four as one surface is the most common mistake in AI visibility work.
TL;DR
- Goodie's GA4 brand panel put AI referral share in March to April 2026 at ChatGPT 62.6%, Claude 18.5%, Gemini 10.6%, Perplexity 7.3%, Copilot 4.0%, against ChatGPT at 89.1% a year earlier.
- Platform size does not predict referrals. In that panel Claude held 1.29% of platform visits and produced 18.0% of referrals. Gemini held 29.0% of visits and produced 10.3%.
- Each engine uses a different crawler for retrieval. Being reachable by one does not mean being reachable by another.
- Google AI Overviews eligibility is governed by ordinary Googlebot indexing. If a page is blocked from Search, it cannot appear as a supporting link.
- AI-sourced sessions in that panel averaged 58.5 seconds of engagement against 44.2 for Google organic, roughly 30% longer.
Referral share, and why it is not the whole picture
Goodie's 2026 AI Search Traffic Report tracked AI referrals across a panel of anonymised B2B sites in GA4 between August 2025 and May 2026. The headline movement:
| Engine | Share, May to Aug 2025 | Share, Mar to Apr 2026 | Change |
|---|---|---|---|
| ChatGPT | 89.1% | 62.6% | -26.5 points |
| Claude | 1.4% | 18.5% | +17.1 points |
| Gemini | 2.4% | 10.6% | +8.2 points |
| Perplexity | 3.1% | 7.3% | +4.2 points |
| Copilot | 3.2% | 4.0% | +0.8 points |
Two cautions before anyone reallocates a budget on this. It is a B2B panel, so it says nothing about consumer categories. And referral share measures only the traffic that arrives with an identifiable source, which by construction excludes Google AI Overviews, where the answer usually ends the session.
The more useful figure in that report is the mismatch between platform size and referral output. Claude held 1.29% of measured platform visits and produced 18.0% of B2B referrals. Gemini held 29.0% of visits and produced 10.3%. Grok held 3.5% of visits and produced effectively no B2B referrals at all. Sizing your effort to platform popularity gets this exactly backwards.
The mechanics matrix
| ChatGPT | Claude | Perplexity | Google AI Overviews | |
|---|---|---|---|---|
| Retrieval crawler | OAI-SearchBot | Claude-SearchBot | PerplexityBot | Googlebot |
| Live fetch agent | ChatGPT-User | Claude-User | Perplexity-User | None |
| Searches by default | Sometimes. Model decides | Sometimes. Model decides | Always | Always |
| Citation posture | Layered onto the answer | Core to the answer pattern | Citation-first by design | Supporting links, low prominence |
| Typical query stage | Broad discovery | Deeper research | Source verification | Informational, top of funnel |
| Sends clicks | Yes, moderately | Yes, disproportionately | Yes, by design | Rarely |
| Honours robots.txt | Yes | Yes | Documented as ignored for Perplexity-User | Yes |
ChatGPT
ChatGPT does not search every query. The model decides whether a question needs live retrieval, and for a large share of prompts it answers from parametric knowledge instead. That single behaviour splits your work in two: appearing in the search index via OAI-SearchBot governs the retrieved answers, and being described accurately in the training data governs everything else. GPTBot is the training crawler and OAI-SearchBot is the retrieval one, and blocking the first does not affect the second.
Practical consequence: for a category where ChatGPT usually answers without searching, structural optimisation moves very little. What moves it is being written about across enough independent sources that the model's underlying knowledge of your category names you. That is a public relations and community problem more than a content-structure problem.
Highest-leverage action: confirm OAI-SearchBot has access, then invest in third-party mentions rather than on-page changes.
Claude
Claude is the engine most B2B teams underweight, and the Goodie panel is the clearest argument against doing so. Fourteen of sixteen brands with sufficient Claude volume moved upward between waves, which the report tested at a binomial p-value of 0.0021.
Claude's usage skews toward longer research-shaped questions, and its answers cite sources as a matter of default pattern rather than as an add-on. Longer, more specific questions favour content that goes deep on a narrow thing. A 1,200-word page that fully answers one technical question tends to do better here than a 5,000-word category overview.
Anthropic runs three agents: ClaudeBot for training, Claude-SearchBot for retrieval, Claude-User for live fetches. All three are separately controllable in robots.txt, and the retrieval one is the one that produces citations.
Highest-leverage action: allow Claude-SearchBot, then publish narrow, technically complete pages rather than broad ones.
Perplexity
Perplexity always retrieves. There is no parametric-answer path to compete with, which makes it the most legible engine to optimise for and the best diagnostic surface. If your content is structurally sound and you are absent from Perplexity, the problem is access or authority, not structure.
Perplexity documents PerplexityBot as an indexing crawler that is explicitly not used for foundation model training, which removes the usual objection from legal. Perplexity-User handles user-initiated fetches and is documented as generally ignoring robots.txt, so blocking it requires an edge rule.
Because Perplexity is citation-first, it is where clean source attribution pays most directly. Pages with named sources, dates, and specific numbers surface here more readily than pages making the same claims without attribution.
Highest-leverage action: use Perplexity as your weekly test surface. It gives the fastest feedback of the four.
Google AI Overviews and AI Mode
Google is the largest surface by audience and the least generous with clicks. AI Mode's zero-click rate runs around 93%.
The mechanics are the most conventional of the four. Eligibility is governed by ordinary Search indexing: a page blocked from Google Search cannot appear as a supporting link in AI Overviews or AI Mode. Google-Extended, despite the name, is not a crawler at all. It is a control token governing whether your content trains Gemini and whether it is used for grounding at prompt time, and disallowing it does not affect Search ranking or AI Overviews eligibility.
Google also decomposes queries before retrieving, which it calls query fan-out. Optimising a single page for a single head term leaves most of the fanned-out sub-questions unanswered.
Highest-leverage action: ordinary technical SEO, plus coverage of the sub-questions rather than only the head term. There is no separate AI Overviews optimisation lever.
What the engagement data says
In the Goodie panel's February to May 2026 window, AI sources averaged 58.5 seconds of engagement with a 67.8% engagement rate, against Google organic at 44.2 seconds and 61.8%, and Bing organic at 47.9 seconds and 58.4%. Roughly 30% longer than Google organic.
The mechanism is plausible: someone arriving from an AI answer has already been told what your page contains and why it is relevant, so the mismatch between expectation and content is smaller than for a search result clicked on a headline. It also means a bounce-rate target calibrated on organic traffic will misjudge this channel.
The measurement trap
All of the above understates AI's contribution, for two structural reasons the Goodie report labels the dark traffic problem.
First, a meaningful share of AI-influenced visits arrive as direct traffic. Users copy a URL out of an answer, or read the answer and type the domain later. Native mobile apps frequently strip the referrer entirely.
Second, Google's AI surfaces are bundled into organic attribution. A visit from an AI Overview and a visit from a blue link look the same in Search Console.
So any dashboard built purely on referral source will report a number that is directionally right and materially low. The correction is not a better attribution model, which is not available, but a second measurement axis: a fixed prompt panel run on a schedule, recording appearance and position per engine. That is the only way to see the citations that produce no session at all. Free tools are sufficient to start.
How to allocate effort
| If your priority is | Focus on | Because |
|---|---|---|
| Fastest feedback loop | Perplexity | Always retrieves, so changes show up quickly |
| B2B pipeline per unit of effort | Claude | Overperforms its user base on B2B referrals |
| Raw reach | ChatGPT | Still the majority of measured AI referrals |
| Top-of-funnel category presence | Google AI Overviews | Largest audience, even with few clicks |
| Diagnosing why nothing works | Perplexity, then crawler logs | Isolates structure from access |
For most cybersecurity vendors the order is Perplexity for testing, Claude for pipeline, Google for reach, ChatGPT for the long game of being in the training data. That ordering will look wrong to anyone sizing by user counts, which is the point.
None of this matters if the crawlers cannot reach you, so start with the crawler reference, and none of it survives extraction unless your pages are structured for it, which is a separate discipline. For how buyers actually move between these surfaces, see the enterprise buying view.
Frequently Asked Questions
Which AI engine sends the most B2B traffic?
ChatGPT, at 62.6% of AI referrals in Goodie's B2B GA4 panel for March to April 2026, down from 89.1% a year earlier. Claude was second at 18.5%, having grown from 1.4%. The concentration is falling fast enough that a single-engine strategy is now a measurable risk.
Why does Claude send more traffic than its user base suggests?
Claude cites sources as a default part of its answer pattern rather than as an add-on, and its usage skews toward longer research-shaped questions where a source link is genuinely useful. In the Goodie panel Claude held 1.29% of platform visits and produced 18.0% of B2B referrals, roughly 14 times its visit share.
How do I optimise specifically for Google AI Overviews?
There is no separate lever. Eligibility is governed by ordinary Google Search indexing, so a page blocked from Search cannot appear as a supporting link. The productive work is conventional technical SEO plus covering the sub-questions a query decomposes into, since Google fans a prompt out into multiple retrievals before assembling an answer.
Does blocking GPTBot remove me from ChatGPT search?
No. GPTBot collects training data. ChatGPT's search feature relies on OAI-SearchBot, and live page reads use ChatGPT-User. Blocking GPTBot while allowing the other two keeps you eligible for retrieved answers while opting out of training.
Which engine should I test with?
Perplexity. It always retrieves rather than sometimes answering from memory, so a structural change to a page produces observable feedback quickly. If your content is well structured and you are still absent from Perplexity, the problem is crawler access or authority rather than the writing.
Is AI traffic more engaged than search traffic?
In the Goodie panel, yes: AI sources averaged 58.5 seconds of engagement and a 67.8% engagement rate, against 44.2 seconds and 61.8% for Google organic. Visitors arriving from an AI answer already know what the page contains, which reduces the expectation mismatch that drives quick exits.
Why is my AI traffic underreported?
Two structural reasons. Users copy URLs out of answers or type the domain later, which lands as direct traffic, and native mobile apps often strip the referrer. Separately, Google's AI surfaces are bundled into organic attribution, so an AI Overview visit is indistinguishable from a blue-link visit in Search Console.
Get the newsletter
New writing on identity, AI security, and building software, delivered when it ships. No tracking pixels, no funnels, unsubscribe with one click.