Measurement · AEO + GEO · 22 min read · last updated 2026-09-18
The AEO and GEO experimentation roadmap: from AI visibility to conversions
Ten numbered experiments, the measurement chain from citation to pipeline, and an honest account of where attribution breaks
An AEO and GEO experimentation roadmap is a sequenced programme of numbered tests that moves a brand from invisible in AI answers to cited, and from cited to converting. You fix a bank of representative prompts, sample each engine on a schedule, run one change at a time, and read the result against a measured noise floor. Separately, you instrument the chain that runs from a citation to a referral session to a qualified visit to pipeline, and you write down which links in that chain are measurable, which are only inferable, and which are broken.
Most published AEO advice stops at visibility. That is the wrong place to stop, because visibility is an input and a marketer is asked about output. This guide covers both halves: the experiment programme in the middle, and the measurement chain that decides whether any of it earned anything.
The author of this guide founded GrackerAI, an AI visibility platform, so the measurement problems described here are ones encountered first-hand rather than read about. No vendor is recommended below and no client results are cited.
Why experimentation is different in AI search
Traditional rank tracking assumes a stable ground truth: query a keyword, read a ranked list, and the list is close to identical if you re-run it minutes later. AI search breaks that assumption. The same prompt sent to the same model twice can return a different set of recommended brands, in a different order, with different citations.
SparkToro's research, reported by Search Engine Land, put a number on this. There is "less than a 1 in 100 chance that any AI tool, if asked the same question 100 times, will give the same list of brands in any two responses" (searchengineland.com). That is the central design constraint. If two runs of one prompt rarely agree, one measurement is not a measurement. It is a sample of size one from a noisy distribution.
Three consequences follow:
- Sample, do not spot-check. You need many prompts and repeat runs so the signal averages out. Coverage of intents beats volume on any single query.
- Track distributions, not positions. The useful unit is share across a query set over time, not a rank on a single query.
- Expect drift. Month-over-month citation drift of 40 to 60 percent is reported as normal (authoritytech.io). A 20 percent swing on one query sits inside the noise band. Do not treat it as a result.
The macro backdrop makes the effort worth it. Similarweb data reported by Search Engine Land found 68 percent of US Google searches ended without a click in the first four months of 2026, up from 60.45 percent in 2024. The same reporting found AI Overviews appearing on more than 20 percent of searches and cutting click-through roughly 60 percent (searchengineland.com). Visibility inside the answer, not just the blue link below it, is now the thing to measure.
The measurement chain: citation to conversion
Before any experiment, decide what you can actually observe. The chain has five links. Each one has a different instrument and a different honesty level, and confusing them is how AEO programmes end up reporting numbers that do not survive a finance review.
| Link | What it is | Instrument | Honesty level |
|---|---|---|---|
| 1. Answer impression | Your page is shown inside an AI answer | Search Console generative AI report (Google surfaces only); nothing for assistants | Partly measured for Google, unmeasured elsewhere |
| 2. Citation | Your domain is one of the cited sources | Prompt-bank sampling across engines | Sampled, never a census |
| 3. Referral session | A person clicks through and lands on you | GA4 AI Assistant channel, UTM parameters, server logs | Measured but systematically undercounted |
| 4. Qualified visit | That session shows buying behaviour | Your own analytics: engaged sessions, key events, depth | Fully measured, it is your own property |
| 5. Pipeline | An opportunity and revenue result | CRM, plus self-reported source on the form | Measured only if you ask the buyer directly |
Link 1: answer impressions
For Google surfaces there is now a real instrument. Google's generative AI performance report in Search Console covers AI Overviews and AI Mode, and it reports impressions only. Google's own documentation states it excludes clicks, click-through rate, query data, and Search Labs experiments (support.google.com). Google rolled the report out to all sites worldwide on 31 August 2026.
The older behaviour still applies to the main Performance report. Google documents that sites appearing in AI features "are included in the overall search traffic in Search Console" and are "reported on in the Performance report, within the 'Web' search type" (developers.google.com). So AI Mode clicks are mixed into your organic totals and cannot be separated there.
For ChatGPT, Claude, Perplexity, Copilot and Gemini's assistant surfaces there is no impression report at all. Nobody publishes how often you were shown. Your prompt bank is a survey of a distribution you chose, not a census of what real users asked. Say that out loud whenever you present a presence rate.
Link 2: citations
A citation is observable by sampling: run the bank, record who was cited, compute share. This is the most-instrumented part of the chain and the one every vendor tool sells. Two cautions. First, cross-vendor share numbers are not comparable, because each tool samples a different prompt set. Second, your own share is a property of your bank. Change the bank and the number moves without the world moving.
Link 3: referral sessions
This is where most reporting quietly goes wrong. Three facts matter.
GA4 now has a native AI Assistant default channel. Google defines it as "the channel by which users arrive at your site from sources like ChatGPT, Gemini, Deepseek, Copilot, or Grok," matched when the medium is set to ai-assistant (support.google.com). Crucially, the same documentation notes it excludes Google's AI Overviews and AI Mode, which stay classified under Organic Search. So your AI Assistant channel is assistant traffic only, and your Google AI traffic is invisible inside organic.
ChatGPT appends utm_source=chatgpt.com to outbound citation links, which is why a share of that traffic is attributable at all, and why the parameter travels with a link even when someone pastes it into a document (ppc.land).
Against that, a large share of genuine AI referrals arrives with no referrer at all. Mobile apps that hand a link to the system browser, links opened with a noreferrer attribute, and users who copy a URL out of an answer and paste it into a new tab all land in Direct. Published estimates of the size of this leak vary widely by measurement method and none of them is independently audited, so treat any single percentage as vendor-reported. The direction is not in doubt: your AI referral number is a floor, not a count.
Link 4: qualified visits
This link is the honest one. Once a session is on your property you own the measurement completely. Read engaged session rate, scroll depth, pages per session, and key events for the AI Assistant channel against the same pages' organic numbers. If the AI cohort is too small for stable rates, that is itself the finding, and it tells you to stop optimising the channel and go back to link 2.
Link 5: pipeline
Two instruments, and you need both. The CRM gives you a modelled number from whatever last-touch or fractional model you run, and it will undercount AI badly for the reasons in link 3. A free-text "how did you hear about us" field on your highest-intent form gives you a self-reported number. That is the only instrument that catches the buyer who read an answer, never clicked, and typed your name into a browser a week later. Self-reporting is biased and lossy. It is also the only window onto the people the analytics stack cannot see, so read its direction over quarters rather than its level in any month.
Where attribution genuinely breaks
Be blunt with stakeholders about this. The answer that is read and never clicked is not measurable at the source. No engine reports it, no analytics package receives it, and no amount of tag configuration recovers it. A buyer can read your comparison table inside an AI answer, form a shortlist, and arrive three weeks later as a direct visit or a branded search.
There are three partial proxies and each one is weak on its own:
- Branded query volume in Search Console, lagged 60 to 180 days behind visibility work.
- Direct and branded organic sessions, which move for many reasons other than AI.
- Self-reported source on forms, which catches intent but not volume.
Read all three together, over quarters, as a direction rather than a measurement. Anyone selling you a precise number for the no-click surface is selling a model, not an observation.
What the conversion evidence actually says
The claim that AI referrals convert better than organic is plausible and widely repeated, and the underlying evidence is thinner than the confidence around it.
The most transparent public data point is first-party. Ahrefs reported that AI search accounted for 0.5 percent of its traffic while producing 12.1 percent of its signups over a 30-day window, measured in its own analytics product (ahrefs.com). That is one B2B SaaS company reporting on itself, with a high-consideration signup, and the author flags uncertainty about whether the rate would hold at larger volumes. It is a useful existence proof, not a benchmark.
On volume, Conductor's benchmark analysis of 13,770 domains and 3.5 million prompts put AI referral at 1.08 percent of all traffic and ChatGPT at 87.4 percent of AI referral traffic, drawn from May to September 2025 data (conductor.com). A fast-growing sliver, not a flood.
Other multiples circulate widely: conversion near 14 percent against roughly 3 percent for Google organic, B2B multiples of 4x to 20x, one platform citing 4.4x. Treat every one of these as vendor-sourced and unverified. Use them to justify instrumenting the channel. Never use them to forecast pipeline. The defensible position for a marketer is that the channel is small today, plausibly high-intent, and worth measuring precisely so you can compute your own rate instead of borrowing someone else's.
Building the prompt bank
The prompt bank is the instrument. Everything downstream reads off it, so it has to be representative and stable. Build 40 to 120 prompts per language, prioritising coverage of intent over raw volume (ailabsaudit.com).
Spread the bank across six intent types:
- Branded: prompts that name your company or product directly.
- Non-branded: category questions where you want to surface without being named.
- Comparative: "X vs Y" and "alternatives to Z" prompts.
- Problem-led: the underlying pain, phrased as a user would say it.
- Persona-led: the same problems voiced by distinct buyers, such as a CISO, a founder, a developer.
- Geographic: locale-specific variants where market and language matter.
A practical way to generate the bank: write a one-page scoping note covering your category, top competitors, core problems, target personas and priority geographies, then have an LLM expand it into candidate prompts across the six intents. Edit by hand afterwards. The model gets you breadth fast; your judgment removes the prompts no real buyer would type.
Run two passes per model, because the two modes answer from different places:
- Native pass: browsing off, the model answers from its weights. This shows how the model remembers your category.
- Web pass: the model retrieves and grounds against live sources. This shows which pages get cited right now.
The gap between the passes is itself a finding. Strong native presence with weak web citation means the model knows you but is not grounding on your current pages. The reverse means your pages are citable but the model's baseline knowledge does not include you.
Record the model version with every run. Model releases move baselines, and a test that straddles one is unreadable.
The model set to test
Presence does not transfer between engines. A brand cited heavily by Perplexity can be invisible in Gemini. The core set is ChatGPT, Claude, Perplexity and Gemini. Add Grok, Copilot and Mistral where your geography or audience warrants it.
Weight effort by where your buyers are. Conductor's finding that ChatGPT drives 87.4 percent of AI referral traffic argues for putting the most measurement attention there, while the others still shape perception even when they send less traffic.
The KPI set
Pick a small, honest KPI set and hold it stable across runs. The visibility headline is AI Share of Voice; the business headline is the key-event rate of the AI Assistant channel against your organic baseline.
AI Share of Voice % = (sum of weighted scores for your brand) / (sum of all brands' weighted scores), computed over your query set, across engines, over a time window.
The weighted score is where you encode what matters to you. A citation, a recommendation, and a top-of-list mention can each carry a different weight. What matters is that the weighting is written down and applied identically every run, so movement reflects the world rather than a change in your own accounting.
| KPI | What it answers | Chain link | How to read it |
|---|---|---|---|
| Presence / mention rate | On what percent of prompts do you appear at all? | 2 | Coverage, before quality |
| Citation share | Of all sources cited on your bank, what fraction are yours? | 2 | Rises as pages become groundable |
| AI Share of Voice % | Weighted presence against all brands | 2 | The headline visibility trend |
| Cross-engine consensus | On how many engines do you appear for one intent? | 2 | Durability across the field |
| Sentiment and accuracy | Are you described correctly and positively? | 2 | Quality of presence |
| Generative AI impressions | How often are you shown in AI Overviews and AI Mode? | 1 | Google surfaces only, impressions only |
| AI Assistant sessions | How many people actually arrive? | 3 | A floor, never a count |
| AI key-event rate | Do those sessions do the thing you want? | 4 | Compare against the organic rate, not zero |
| Self-reported AI source | Do buyers name an assistant unprompted? | 5 | Direction over quarters, not level |
Aleyda Solis's three-layer framework is a useful way to organise the set: Presence (do you show up, accurately), Readiness (are the structures in place to be cited), and Business Impact (is value created) (learningaisearch.com).
Cadence
Match measurement frequency to how fast each layer moves. Over-sampling burns budget on noise; under-sampling misses drift.
| Cadence | Scope | Purpose |
|---|---|---|
| Weekly | 10 to 15 priority queries | Early warning on high-value intents |
| Monthly | Full prompt bank sweep | The trend line of record |
| Monthly | Chain links 3 to 5 review | Sessions, key events, self-reported source |
| Quarterly | Strategic review | Reset the bank, competitors, and bets |
Read the monthly full sweep as your source of truth. Because citation drift runs 40 to 60 percent month over month, never draw a conclusion from one week's wobble on one query. The weekly cut is a smoke alarm, not a scoreboard.
The experiment programme
Ten numbered experiments in three phases. Run them roughly in order, one variable at a time, and write the hypothesis down before you look at the result. Retrofitting a story onto a random swing is the easiest way to fool yourself in a non-deterministic system.
| # | Experiment | Phase | Primary metric | Read window |
|---|---|---|---|---|
| E1 | Fetcher access audit | Foundation | Fetch success rate by user agent | 14 to 30 days |
| E2 | Answer-first opening passage | Foundation | Presence rate on matched prompts | 30 to 60 days |
| E3 | Entity and authorship clarity | Foundation | Branded-prompt accuracy | 60 to 90 days |
| E4 | Symmetric comparison tables | Selection | Citation share on comparative prompts | 30 to 60 days |
| E5 | Real review dates | Selection | Citation share on recency-modified prompts | 30 days |
| E6 | Fan-out decomposition | Selection | Sub-question coverage, AI impressions | 45 to 90 days |
| E7 | Original first-party data | Selection | Citations on the specific claim | 60 to 180 days |
| E8 | Next-step blocks on AI entry pages | Conversion | AI key-event rate vs organic | 200+ sessions per variant |
| E9 | Self-reported source capture | Conversion | Share of forms naming an assistant | 60 to 90 days |
| E10 | Brand-demand lift observation | Conversion | Branded impressions, direct sessions | 90 to 180 days |
Phase one: foundation, run these when presence is near zero
E1. Fetcher access audit
- Hypothesis. Engines cannot cite what their fetchers cannot retrieve, so some of your absence is an access problem rather than a content problem.
- Change. Allow the answer-engine fetchers you want in robots.txt, confirm 200 responses per user agent in server logs, and make sure target pages render their substance without client-side JavaScript.
- Metric. Fetch success rate per user agent in logs, then presence rate on the bank.
- Read window. 14 to 30 days for fetch logs, longer for presence to follow.
- Confound. A page can be perfectly fetchable and still uncited on quality grounds. A flat presence rate does not mean the access fix failed. See should you block AI crawlers for the strategic side of this decision.
E2. Answer-first opening passage
- Hypothesis. A direct, unhedged 40 to 60 word answer in the first paragraph is what an engine lifts, and a scene-setting introduction is not.
- Change. Rewrite the opening of ten target pages so the first paragraph answers the page's question outright. Freeze everything else on those pages.
- Metric. Presence rate and citation share on the prompt subset those pages target.
- Read window. 30 to 60 days on tail intents. Do not judge head terms this early.
- Confound. Teams almost always change more than the opener while they are in the file. If the rest of the page moved, the test is void. See citation-worthy content patterns.
E3. Entity and authorship clarity
- Hypothesis. Models describe you inaccurately because your entity is ambiguous, not because they dislike you.
- Change. Ship Organization and Person schema with sameAs links, standardise how the company and product are named across every property, and publish an about page and a methodology page that state plainly what you are.
- Metric. Accuracy of branded-prompt answers across engines, plus the rate of factual errors about you.
- Read window. 60 to 90 days, and longer on the native pass, because parametric memory only updates with model releases.
- Confound. A model release mid-test resets the baseline. This is why you log model versions. See entity authority for AI engines.
Phase two: selection, run these once you are retrievable
E4. Symmetric comparison tables
- Hypothesis. Engines answering comparative prompts prefer pages that compare options on the same attributes, because a symmetric table is trivially extractable.
- Change. Add a real comparison table to one page cluster, with identical attributes for every option and no missing cells.
- Metric. Citation share across the whole comparative subset of the bank.
- Read window. 30 to 60 days.
- Confound. Comparative prompts are the noisiest category in the bank. Read the subset in aggregate, never a single prompt.
E5. Real review dates
- Hypothesis. Prompts carrying a year or a recency word favour pages that show a recent, genuine review date.
- Change. Review a page set for real, correct what is stale, and surface a visible last-reviewed date backed by dateModified.
- Metric. Citation share on the recency-modified subset.
- Read window. 30 days, the fastest read in the programme.
- Confound. Bumping a date without changing the content is a quality risk with no upside, and it makes the test unreadable. Only read this test if the review was real.
E6. Fan-out decomposition
- Hypothesis. Google AI Mode issues many sub-queries per question, so one umbrella page loses to a set of pages that each answer a constituent question cleanly.
- Change. Take one umbrella topic, enumerate the six to ten sub-questions behind it, and make sure each has a page or a dedicated H2 that answers it directly.
- Metric. Sub-question coverage on the bank, plus generative AI impressions for the page set in Search Console.
- Read window. 45 to 90 days.
- Confound. You also created new URLs, so total impressions rise mechanically. Judge per-question coverage, not the sum. See how Google AI Mode works.
E7. Original first-party data
- Hypothesis. A number nobody else has is the most citable thing you can publish, because a synthesiser has no substitute source for it.
- Change. Publish one genuine first-party dataset, survey, or audit, with the method stated and the raw counts shown.
- Metric. Citations pointing specifically at the new claim, plus third-party pickup of the number.
- Read window. 60 to 180 days.
- Confound. Original data also attracts links and press, so the effect is entangled with classical authority. That entanglement is real and does not need to be separated.
Phase three: conversion, run these once you are cited
E8. Next-step blocks on AI entry pages
- Hypothesis. Someone arriving from an AI answer has already done the research and needs the next step, not the introduction they just read a summary of.
- Change. Add an explicit next-step block near the top of your ten highest AI-referred entry pages.
- Metric. Key-event rate for the AI Assistant channel on those pages, compared against the same pages' organic rate.
- Read window. Measured in sessions, not days. Wait for roughly 200 AI-referred sessions per variant.
- Confound. AI referral volume is small, so underpowered tests are the default failure mode here. If you cannot reach the session count in a quarter, you do not have a conversion problem yet. Go back to phase two.
E9. Self-reported source capture
- Hypothesis. A share of your pipeline already comes from AI answers and is currently booked as Direct.
- Change. Add a free-text "how did you hear about us" field to your highest-intent form, and tag responses that name an assistant.
- Metric. Share of inbound naming an assistant by name.
- Read window. 60 to 90 days, or 100 responses, whichever comes later.
- Confound. People under-report and misremember their own path. Read the direction across quarters, never the level in one month.
E10. Brand-demand lift observation
- Hypothesis. Answers that are read and never clicked show up later as branded searches and direct visits.
- Change. None on site. This is an observation test riding on the phase one and two work.
- Metric. Branded query impressions in Search Console and direct session volume, both lagged behind the content work.
- Read window. 90 to 180 days.
- Confound. Every other thing marketing does moves these numbers too. This is the weakest measurement in the programme, and it should be labelled as such in every report where it appears.
Reading a result honestly
Three rules keep the programme from producing fiction.
Compare against your own drift band, not against zero. Measure your bank's month-over-month variance during the baseline period and write the number down. A change that moves a metric by less than that band has shown nothing, however pleasing the direction.
Respect the read window. Long-tail queries can respond in 30 to 60 days; competitive head terms take 3 to 6 months (authoritytech.io). Judging a head-term test at three weeks guarantees a false read.
Do not claim causation you cannot support. There is no rigorous, peer-reviewed protocol proving that a specific formatting change causes a citation lift. You also cannot run a clean holdout, because serving engines a different version of a URL than you serve people is cloaking. The honest framing is that you are steadily improving groundable pages and watching a noisy aggregate trend upward over months.
What to do first, and when to stop
The decision rule depends on where you are, not on what is fashionable.
| Your situation | Run | Do not run |
|---|---|---|
| Presence under about 10 percent of the bank | E1, E2, E3 | Any conversion work. You have a retrieval problem, not a conversion problem |
| Present but low citation share on comparative prompts | E4, E5, E6 | More net-new pages on topics you already cover |
| Cited well, few AI sessions | E7, and check link 3 instrumentation before concluding anything | Rewriting content. Your measurement may be the problem, not your pages |
| AI sessions arriving, key-event rate below organic | E8, E9 | Adding pages. Fix the landing experience first |
| Cited, converting, channel still around 1 percent of sessions | Hold the loop, spend the marginal hour elsewhere | Scaling investment on a channel-size assumption you have not verified |
When to stop a test. Stop when the effect is inside your measured drift band after two full monthly sweeps, or when the read window has passed with no directional movement. Revert the change only if it cost something; a neutral improvement to a page is still an improvement.
When to stop expanding the programme. Stop growing the prompt bank when new prompts stop changing the aggregate. Stop adding engines when an engine's referral contribution stays below the effort of sampling it. A bank of 400 prompts across eight engines that nobody reads is worse than 60 prompts across four engines that drive a monthly decision.
What has no reliable measurement yet
The discipline is young, and pretending otherwise is how teams lose credibility. These are genuinely open:
- No controlled causal test exists for citation lift. No holdout, no clean A/B, and no peer-reviewed protocol. Everything in the programme above is a before-and-after observation against a noise floor.
- Assistant impressions are unpublished. You cannot know how often you were shown inside ChatGPT, Claude or Perplexity. Prompt-bank presence is a proxy for a number nobody reports.
- Google reports AI impressions without clicks. The generative AI performance report carries impressions only, with no clicks, CTR or query data, and AI Mode clicks stay merged into the Web search type.
- The no-click answer is unmeasured at source. The largest AI surface by volume produces no telemetry for publishers at all.
- Cross-vendor visibility numbers are not comparable. Each platform samples a different prompt set, so two tools can report different shares for the same brand and both be correct.
- Conversion multiples are not benchmarks. The public figures are vendor-sourced or single-company first-party reports. Compute your own.
Anyone who tells you these are solved is describing a product roadmap rather than the current state of measurement.
A 90-day rollout
Days 1 to 30: instrument.
- Write the one-page scoping note covering category, competitors, problems, personas and geographies.
- Generate 40 to 120 prompts across the six intents, then hand-edit down to a stable bank.
- Fix the model set and log model versions from the first run onward.
- Write down the AI Share of Voice weighting before the first measurement.
- Run the baseline: native and web passes per model, across the whole bank.
- Instrument chain links 3 to 5: confirm the GA4 AI Assistant channel is reporting, separate Google AI surfaces from assistant traffic in your reporting, connect the Search Console generative AI report, and add the self-reported source field (E9).
Days 31 to 60: establish rhythm and start testing.
- Start the weekly cut on 10 to 15 priority queries.
- Run the second monthly sweep and record the drift band you observe on your bank. That number is now your significance threshold.
- Ship E1, then E2 on tail intents where a 30 to 60 day read is realistic.
- Identify the highest-value intents where you are absent or under-cited.
Days 61 to 90: read, connect, decide.
- Run the third monthly sweep. Three points let you separate trend from noise.
- Judge E1 and E2 against the drift band, not against zero.
- Report the full chain to stakeholders, with each link labelled measured, sampled or inferred.
- Hold the quarterly review: retire stale prompts, add competitors, re-scope the bank, and pick the next two experiments from the decision table.
From there the loop repeats: weekly smoke checks, monthly truth, quarterly re-scope, one honest test at a time. For the tooling that automates the tracking, see choosing an AI visibility tool. For the metric definitions in depth, see measuring AI visibility. For how the two disciplines differ, see AEO vs GEO explained.
Last verified: September 2026. Sources checked for this update: Google Search Console generative AI performance report documentation, Google Analytics default channel group documentation, Google Search Central AI features documentation, Ahrefs first-party conversion report, and Conductor's AEO and GEO benchmark report.
Related guides
- Measuring AI visibility: KPIs, instrumentation, and what to actually track
- How to choose an AI visibility tool: a buyer's guide to the GEO tooling category
- Citation-worthy content patterns: writing for both extraction and grounding
- AEO vs GEO: how Answer Engine Optimization and Generative Engine Optimization actually differ