Skip to content

How to Audit What ChatGPT, Perplexity, and Google Say About Your Brand

GEO/AEO · practitioner · 11 min read · last reviewed 2026-08-20

The full protocol for auditing AI search visibility: build a 60 to 150 prompt universe, test search on and off across engines, read the diagnosis matrix, and build a source ledger that becomes your roadmap.

TL;DR

  • Build 60 to 150 prompts across eight categories (identity, category, recommendation, comparison, objection, jobs-to-be-done, integration, pricing), sourced from real sales calls and search queries, not guesswork.
  • Run every prompt three times per engine, with web search on and off separately; the on/off split tells you whether you have a training problem or a retrieval problem.
  • Cross search-off against search-on to get four diagnoses: healthy, retrieval failure, retrieval-dependent, or invisible, each with a different fix and timeline.
  • Build a source ledger of every citation, sorted by frequency; the top 20 cited domains usually explain most of what the machine believes about your category.
  • 51% of B2B buyers now start research in an AI chatbot (G2, March 2026), so a stale or absent audit is a live gap in the buying process, not a future risk.
The AI visibility diagnosis matrix A two by two grid crossing whether a brand appears with web search on against whether it appears with web search off. Present in both is healthy. Present with search off only means retrieval failure. Present with search on only means retrieval dependent. Absent in both means invisible, requiring a six to twelve month build. Search on: you appear Search on: you are absent Search off: you appear Search off: you are absent HealthyDefend and expand Retrieval failureFix crawling and structure Retrieval dependentDeepen third-party proof InvisibleFull build, 6 to 12 months
Figure 3. Cross the search-off result against the search-on result. Each quadrant needs a different fix on a different timeline.

An AI visibility audit tests what ChatGPT, Google AI Overviews, Perplexity, and Claude actually say about your brand, run across 60 to 150 real buyer prompts, three times each, with web search on and off. The output is a diagnosis matrix that tells you whether you have a training problem, a retrieval problem, or both, plus a ranked source ledger of exactly which pages the machine is reading. Five prompts and a vibe check produce noise. This produces a roadmap.

I run GrackerAI, so I will say the commercial part out loud: we build tooling around exactly this workflow. The protocol below is the same one you would run by hand with a spreadsheet, three browser tabs, and a free afternoon. The tooling saves time at scale, it does not change the method.

The urgency is real. G2's March 2026 survey of 1,076 B2B software buyers found 51% now start research in an AI chatbot rather than a search engine, up from 29% in April 2025. 69% said chatbot guidance changed their vendor selection. Forrester's 2026 survey of roughly 18,000 business buyers found 94% used AI during their most recent purchase. An answer that never names you is a shortlist you never make.

AI referral traffic itself is still small in absolute terms. Cloudflare Radar regularly shows Google sending the overwhelming majority of search referrals while all AI chatbots combined send a fraction of a percent. Do not build the business case for this audit on session counts. Build it on answer share: whether you are named, not whether anyone clicked.

Build the prompt universe

Do not test five prompts. Build 60 to 150 across eight categories, and pull the actual wording from real sources, not your imagination.

CategoryCountExample (cloud security vendor)
Identity10 to 15What is [Brand]? Who founded it and when?
Category10 to 15What is CNAPP? Who are the leading vendors?
Recommendation15 to 25Best CNAPP for a 500-person fintech
Comparison10 to 20[Brand] vs [Competitor], alternatives to [Competitor]
Objection and risk8 to 12Is [Brand] secure? Has it had a breach?
Jobs-to-be-done10 to 15How do I stop developers committing secrets to Git?
Integration and technical5 to 10Does [Brand] integrate with Splunk?
Pricing and procurement5 to 10How much does [Brand] cost?

Sources ranked by usefulness: sales call transcripts (search for "we asked ChatGPT" and "we looked at"), support tickets, Google Search Console question-shaped queries, Reddit and Slack threads in your category, and win/loss interview notes. Add fan-out capture too: run a head prompt and log the sub-queries the engine actually generates. How query fan-out works explains why those sub-queries are usually better targets than the head prompt itself.

Run the test protocol

Sloppy testing produces noise that looks like insight. Lock the protocol before you run anything.

Isolate the environment: log out or use a fresh account with memory and personalization disabled, start a new chat per prompt since context bleeds badly, clear custom instructions, and record the geography and model version, since answers vary by both.

Run each prompt twice, in two modes. Search off, or a model with no browsing, reads parametric memory, the training-layer signal. Search on reads the retrieval and live-fetch layers. This split is the single most valuable thing in the whole audit; how AI search works is the model behind why it matters.

Run three times per prompt per engine minimum, five for your top 20 commercial prompts. These systems are stochastic, so a single observation tells you nothing. Log whether search actually fired: a prompt that did not trigger retrieval says nothing about your retrieval performance and should be bucketed separately, not counted as a loss.

Record, per run: prompt, engine, model version, date, geography, search on/off, whether the brand was mentioned, its position in the list, competitors mentioned in order, sentiment, an accuracy score from 0 to 3, any specific factual errors verbatim, every source URL cited, and whether search actually fired.

Read the diagnosis matrix

Cross search-off against search-on. Four quadrants, four different problems with four different fixes and timelines.

Present in both is healthy: defend the position and expand into adjacent prompts. Present with search off only, absent with search on, is a retrieval failure: the model knows you but cannot find current evidence, usually a crawler access, indexing, or page-structure problem, and it is the fastest fix in this guide. Absent with search off, present with search on, is retrieval-dependent: you exist only when the engine looks, which is fragile but workable, and the fix is deepening third-party consensus so the next training run absorbs you. Absent in both is invisible: not in training data and not retrievable, requiring a full-stack build across entity records, third-party consensus, earned media, and content, typically 6 to 12 months.

Two more diagnoses hide inside the matrix. Present but misdescribed is when the machine names you and gets you wrong: wrong category, wrong founding year, an acquired company still listed as independent. Treat this as urgent, because a confident wrong answer disqualifies you more actively than absence does. Cited but not recommended is when you sit in the source list and never the recommendation, usually a sign your educational content is strong but your comparison and use-case content is thin.

Build the source ledger

Trace every citation back to its origin and build one row per unique URL, aggregated across all prompts and engines: the domain, source type (owned, earned, UGC, review, reference, aggregator, competitor), which prompts and engines cited it, whether it contains a fact about you, whether that fact is accurate, sentiment, whether you control it, publication date, and a remediation owner and due date.

Sort by citation frequency. The top 20 domains typically account for most of what the machine believes about your category, and that list is your real roadmap, not the content calendar you had planned. Fixing what those specific domains say is worth more than publishing ten more pages on your own site, because most citations come from outside it in the first place.

A worked example, illustrative

To make the matrix concrete: a hypothetical CIAM vendor runs 84 prompts across five engines, three runs each. Identity prompts land at 100% presence everywhere, fine. Category prompts land at 62%, named as a CIAM vendor but never as a leader. Recommendation prompts land at 18%, with two competitors appearing in over 70% of runs. In 6 of 11 comparison prompts, the cited source is a competitor's own comparison page, meaning the competitor's framing of the vendor is the only framing available. Two engines report a false "acquired in 2023" claim, traced to a mis-headlined trade article. Search-off and search-on both show absence on recommendation prompts, the bottom-right quadrant.

The roadmap that falls out of it, in order: kill the acquisition error at the source, publish eight use-case pages targeting the specific recommendation prompts where the vendor is absent, publish comparison pages so the competitor's framing is not the only one available, get listed in the "best CIAM" listicles already driving 60% of cited sources, and push a review campaign on the platform that appeared in 40% of citations. Nothing exotic. The audit made the sequence obvious, which is the point of running one properly.

Run it as a loop, not a one-off

A single audit has a shelf life of a few weeks. Run the full prompt set monthly and your top 20 commercial prompts weekly, on a fixed schedule, with the same prompt set version noted on every report, so the numbers stay comparable month over month.

Set expectations by lever before you start: crawler and access fixes show a first signal in 1 to 7 days and a full effect in 2 to 4 weeks; schema and entity records take 2 to 4 weeks to 6 to 12 weeks; new owned content takes months, not weeks; and a training-data shift, the search-off answer finally changing, runs 6 to 18 months on the vendor's own release schedule. Setting up AI crawler access is the fastest lever on that list, so start there while the audit results are still fresh.

Key takeaways

  • Five prompts is not an audit. Three runs per prompt minimum, because these systems are stochastic and a single observation misleads in both directions.
  • A confident wrong answer about your brand is worse than being ignored; treat 'present but misdescribed' as more urgent than 'absent'.
  • Being named in the source list without being the recommendation usually means your comparison and use-case content is thin, not your educational content.
  • Run the full audit monthly and your top 20 commercial prompts weekly; a single audit has a shelf life of a few weeks.
  • Measure answer share, not AI referral traffic; the traffic number is still small in absolute terms even as it decides shortlists.

Frequently asked questions

How many prompts do I need for an AI visibility audit?
60 to 150 prompts across eight categories: identity, category, recommendation, comparison, objection and risk, jobs-to-be-done, integration, and pricing. Five prompts produces noise, not a baseline.
What does search on vs search off tell you?
Search off reads the model's parametric memory, what it learned in training. Search on reads the retrieval and live-fetch layers. If a brand appears with search off but not search on, that is a retrieval failure, a different fix than a training-data gap.
What is the AI visibility diagnosis matrix?
A two by two grid crossing search-off presence against search-on presence: healthy, retrieval failure, retrieval-dependent, or invisible. Each quadrant needs a different fix on a different timeline.
How often should I re-run the audit?
Full prompt set monthly, top 20 commercial prompts weekly. A one-off audit has a shelf life of a few weeks because engines and their answers change constantly.
How long until an AI visibility fix shows results?
Crawler and access fixes show a first signal in 1 to 7 days. Schema and entity records take 2 to 12 weeks. Training-data shifts, the search-off answer changing, take 6 to 18 months on the model vendor's own release schedule.

Related

← All How-To & Implementation guides