You Can Do Everything Right and Still Be Invisible
Every GEO playbook assumes AI engines can read your site. Cloudflare flipped the default to blocked on new domains and serves roughly a fifth of the web. I look at the access layer nobody optimizes, why it usually sits with infrastructure teams, and how to audit yours.

Every piece of GEO advice you have read assumes something it never states: that AI engines can actually reach your content.
The short answer up front: most sites are readable, so this is not a universal crisis, but a meaningful minority are invisible to AI engines for reasons that have nothing to do with content quality. Cloudflare serves roughly 20% of all web traffic and has blocked AI crawlers by default on new domains signing up to Cloudflare since July 1, 2025. Squarespace ships a crawler-blocking toggle. Getting this wrong makes every other GEO investment worthless, and the people who control the setting usually do not report to marketing.
The access layer nobody optimizes
Start with what changed. Cloudflare spent two years building the control side of the AI relationship before it built the visibility side. AI Crawl Control launched as AI Audit in September 2024 and reached general availability in August 2025. It lets site owners see every AI bot hitting their domain and decide per crawler whether to allow it, block it, or charge it. Pay Per Crawl revived the HTTP 402 status code so crawler access could be negotiated machine to machine. On what Cloudflare called Content Independence Day, July 1, 2025, it flipped the default for new domains joining Cloudflare to block AI crawlers unless the operator opts in.
Then, from September 15, 2026, Cloudflare began blocking mixed-use crawlers on ad-supported pages by default for new customers, new sites of existing customers, and all existing free-tier customers.
Read that list of defaults again. If you added a domain to Cloudflare after mid-2025, or you are on a free tier, your posture toward AI engines was set by someone else's policy decision, and it was set to no.
Give the publishers their due
The skeptical read here is that this is a manufactured problem, and there is a real argument on the other side that I want to state at full strength.
Cloudflare did not do this arbitrarily. AI crawlers scrape pages many times for every referral they send back. Publishers watched their content train and ground systems that returned almost no traffic, and blocking was a rational response to an extractive relationship. If you produce original research or journalism, defaulting to blocked is a defensible business decision, not a misconfiguration.
Cloudflare itself named the tension. By July 2026 it had publicly acknowledged that blocking everything could make smaller sites invisible. It described the choice facing operators as a Faustian bargain: stay discoverable and let AI train on your content, or block it and risk losing discoverability.
The honest scope of the problem also matters. In one test of 744 live sites across major platforms, roughly four out of five arrived fully readable to AI crawlers, and the hosting platform barely moved that number. So this is not the hidden explanation for most brands' AI invisibility. Anyone selling it that way is doing the thing I criticize vendors for elsewhere in this series.
But the minority it affects cannot fix it with content, and most of them do not know they are in it.
The distinction that actually matters: training versus retrieval
Here is where I see competent teams make an expensive mistake.
"AI crawler" is not one thing. GPTBot, CCBot, and ClaudeBot are primarily training crawlers. OAI-SearchBot, ChatGPT-User, PerplexityBot, Claude-SearchBot, and Claude-User are retrieval crawlers that fetch pages to answer live questions. Google-Extended governs Gemini training and grounding separately from Googlebot.
Block the retrieval crawlers and you remove yourself from citation in live AI answers. Block only the training crawlers and you keep your content out of model weights while staying eligible to be cited. Those are completely different business decisions, and a blanket "block AI bots" toggle collapses them into one.
I see the collapsed version constantly. A team decides it does not want its content training models, someone flips a single switch, and eighteen months later the company wonders why it never appears in ChatGPT while its competitors do. The content was fine. The policy was never revisited.
This is an access control problem wearing a marketing costume
The structural reason this persists is organizational, and it is the part that interests me most given where I spend my time.
Crawler policy lives in robots.txt, in CDN configuration, and in WAF rules. Those are owned by infrastructure, platform, or security teams. AI visibility outcomes are measured by marketing. The control and the consequence sit in different reporting lines, which is the classic condition for a gap nobody owns.
Anyone who has worked in identity will find this familiar. It is an authorization problem: which non-human identities may access which resources, under what conditions, with what rate limits, for what purpose. We have spent years building governance for machine identities inside the enterprise. AI crawlers are machine identities operating at the perimeter, making requests with declared intent that may or may not match their behavior. Cloudflare's own product line now includes monitoring whether crawlers comply with the directives they were given.
Treat it accordingly. This is not a marketing setting. It is an access policy with a revenue consequence, and it deserves the same review cadence as any other access policy.
How to audit yours this week
Four checks, in order, none of which require a vendor.
Fetch your own robots.txt and read it rather than assuming. If you are behind a CDN with managed rules, the live file may include directives injected at the edge that are not in your repository.
Separate your crawler list into training and retrieval, then confirm the split matches a decision someone actually made. If nobody can name the person who decided, you do not have a policy, you have a default.
Check your platform-level toggles independently of robots.txt. Squarespace, for example, exposes a block for known AI crawlers in settings, and platform-managed robots files list crawler names you did not add.
Look at server logs for retrieval crawler hits and errors. Access is binary and observable. Either OAI-SearchBot and PerplexityBot are fetching your pages successfully or they are not, and that is a log query, not an inference.
Access is necessary, not sufficient
One caveat I want to be precise about, because this is where the story gets oversold.
Being crawlable does not make you cited. Cloudflare can tell you which AI operators crawled you, which pages they requested, what errors they hit, and how much referral traffic came back. What it cannot see, and what nobody outside the model providers can see, is inside the AI conversation itself. Cloudflare acknowledges this limit openly, which is to its credit.
So the correct mental model is two gates. The first gate is access, which is binary, cheap to verify, and occasionally catastrophic when misconfigured. The second gate is whether the engines choose to recommend you, which is probabilistic, expensive to move, and where the actual GEO work lives.
Most of this series argues about the second gate. This piece is about the one nobody checks.
For B2B software companies, the asymmetry is brutal. My data at GrackerAI (disclosure: my company, a GEO platform for B2B SaaS) shows 4 in 10 B2B security buyers start vendor research inside AI assistants. A blocked retrieval crawler removes you from that entire research path silently, with no error message, no ranking drop, and no dashboard turning red.
Here is my question: do you know, without checking, whether your site allows retrieval crawlers today, and can you name the person who decided? For most companies I ask, the honest answer to both is no, and that is a fifteen-minute fix sitting in front of a year of content investment.
Frequently Asked Questions
Does Cloudflare block AI crawlers by default?
Yes, for some sites. Since July 1, 2025, which Cloudflare called Content Independence Day, new domains joining Cloudflare block AI crawlers by default unless the operator opts in. From September 15, 2026, Cloudflare also began blocking mixed-use crawlers on ad-supported pages by default for new customers, new sites of existing customers, and existing free-tier customers. Cloudflare serves roughly 20% of all web traffic.
What is the difference between AI training crawlers and retrieval crawlers?
Training crawlers such as GPTBot, CCBot, and ClaudeBot collect content used to train models. Retrieval crawlers such as OAI-SearchBot, ChatGPT-User, PerplexityBot, and Claude-SearchBot fetch pages to answer live user questions. Blocking retrieval crawlers removes a site from citation in AI answers, while blocking only training crawlers keeps content out of model training while remaining eligible for citation. A single blanket block collapses two different business decisions into one.
How do I check whether AI engines can access my site?
Fetch your live robots.txt rather than reading the repository version, since CDN-managed rules can inject directives at the edge. Separate your crawler directives into training and retrieval and confirm someone deliberately made that split. Check platform-level toggles separately, such as Squarespace's block for known AI crawlers. Then query server logs for successful fetches and errors from retrieval crawlers, since access is binary and observable.
Does allowing AI crawlers guarantee AI visibility?
No. Access is necessary but not sufficient. Crawl access is binary and verifiable through logs, while being recommended in AI answers is probabilistic and depends on authority, structure, and source relationships. Even Cloudflare, which sits between websites and incoming traffic, acknowledges it cannot measure AI answers the way it measures crawls. Roughly four out of five sites in one 744-site test were already fully readable, so access explains a minority of AI invisibility rather than most of it.
More like this
All GEO & AI Search- GEO & AI Searchllms.txt Is the Meta Keywords of 2026Google says llms.txt neither helps nor harms. John Mueller compared it to the meta keywords tag. Adoption sits near 10% of domains and…
- GEO & AI SearchWho Is Buying and Building AI Visibility: An Acquisition Map for the GEO CategoryAdobe paid $1.9B for Semrush. HubSpot acquired a startup under a year old. Wix, Squarespace, Optimizely, and Cloudflare built natively.…
- GEO & AI Search10 Lessons From Tracking 50,000 AI Citations Across 6 EnginesOver 90 days I tracked how six AI search engines cite sources across 50,000+ B2B software responses. The data broke several assumptions…
Get new GEO & AI Search writing
Enjoyed this? Subscribe and tell us what you read most. GEO & AI Search is already ticked for you. No tracking pixels, unsubscribe with one click.