Skip to content
By Opinion

Publishers Are Ready to Block Google. Blocking Is Not a Strategy.

Reddit, USA Today and Reuters are weighing whether to cut Google off. Most of the numbers from that story were corrected a day later. What the breaking crawl bargain means for B2B SaaS, and why blocking is a negotiating position rather than a plan.

Publishers Are Ready to Block Google. Blocking Is Not a Strategy., by Deepak Gupta on guptadeepak.com

45.2 million impressions. A 3.14% average click-through rate. Average position 6.1.

That is not a wire service's dashboard. That is my own site, and by the standards of this market it is a good chart. Position 6.1 means the rankings are working. It also means 43.8 million of those impressions produced no visit at all, and that is the healthy version of this picture. I have been watching that gap for two years while writing about why it widens.

So when the Wall Street Journal reported on July 21 that Reddit has internally discussed shutting Google out of its content for AI use, and that USA Today, Politico, Reuters, the Economist and People Inc. are all reassessing whether they work with Google at all, my first reaction was not surprise. It was recognition. The publishers have arrived at the number I have been staring at. Reddit's stock fell 9% the next day on the report alone.

My second reaction was more useful: most of the people forwarding that story are drawing exactly the wrong conclusion from it.

The numbers everyone quoted this week were wrong

The original version of the story carried Semrush figures that turned out to be inaccurate. The Journal corrected them on July 22. By then the first set had already been reprinted across a dozen tech sites, and it is still circulating.

Here is the gap between what got shared and what is actually true, comparing U.S. organic Google search traffic in June 2025 to June 2026.

PublisherAs first reportedAs corrected (July 22)
USA Today (national paper)down nearly 50%down 18%
Business Insiderdown more than 85%down 31%
CNNdown about 25%down 31%
Politicodown 23%down 20%

The trailing twelve-month averages in the Journal's own chart were revised too, and that revised set is the more interesting one, because it does not point in a single direction.

PublisherTrailing 12-month change
The Guardianup about 16%
BBCup about 15%
Peopledown 3.2%
The Wall Street Journaldown 18%
Timedown 27%
USA Todaydown 28%
Reutersdown 35%
Business Insiderdown 43%
The Washington Postdown 44%

Three things follow from that.

First, an 18% decline and a 50% decline produce very different board meetings. One is a hard year. The other is a business model ending. Several publishers are being asked to make an irreversible infrastructure decision based on a version of the data that no longer exists.

Second, the spread is the story. If AI Overviews were a single force acting uniformly on the web, the Guardian and the BBC would not be up while the Washington Post is down 44%. Something else is doing most of the work: how much of a catalog is commodity information a model can answer without sending anyone anywhere, how much brand demand exists independent of search, and how well a publisher has moved its audience into apps, newsletters and direct visits.

Third, and this is the part I keep coming back to: the aggregators got it wrong because they were summarizing a summary. That is a live demonstration of what happens to second-hand analysis in an AI-mediated web. It is also the whole argument for publishing something first-hand.

What is actually breaking is the crawl bargain

For twenty years the deal was legible. You let the crawler in, you got traffic back. One bot, one purpose, one exchange rate.

That deal now has one bot and four purposes. The same fetch feeds the search index, grounds AI Overviews, grounds AI Mode, and in some configurations informs training. Publishers can block Google-Extended and keep their material out of Gemini training, but Googlebot still crawls for the AI features inside Search itself. Turning off the answer has generally meant turning off the index, which is not a choice so much as a dare.

Two things are changing that.

On June 3 the U.K. Competition and Markets Authority imposed a conduct requirement that forces Google to let publishers opt out of AI features without losing their place in ordinary search results, and to attribute publisher content properly inside AI-generated answers. Google said it will comply, starting with a U.K. test, and has not committed to a global timeline. That is the single most consequential unpriced variable in this whole story, because it converts an all-or-nothing decision into a dial. If it ships globally, half the leverage publishers are threatening to use this week evaporates, because they will be able to withhold the answer without withholding the index.

Cloudflare is moving faster. From September 15, 2026, mixed-use crawlers get blocked by default on any page carrying ads, for new customers, new sites on existing accounts and free-tier accounts. Cloudflare has split bots into three categories, search, agent and training, and applies the most restrictive rule to any crawler that blends them. Googlebot blends them. The company also replaced Pay Per Crawl with Pay Per Use, which pays when content shows up in an answer rather than when a bot fetches a page. Cloudflare says bots are now more than half of all web traffic, and that mixed-use crawlers account for roughly 36% of crawler activity.

Read that as what it is. A large piece of internet infrastructure is trying to force a price onto something that has been free since 1998, and it is doing it by changing a default rather than by asking. If your site sits behind Cloudflare and you have never opened the bot settings, a decision you did not make is scheduled for September 15.

Blocking is a negotiating position, and most of us do not have one

Reddit can talk about cutting Google off because Reddit has three things at once: content no model can synthesize its way around, a contract reportedly worth $60 million a year, and a renewal date. That combination is what negotiating power looks like. Take away any one of the three and the threat stops being credible.

Before anyone in B2B copies the posture, run the three questions honestly. The answers usually look like this.

QuestionRedditTypical B2B SaaS blog
Can a model answer your queries without you?No. The content is unique human argument.Yes. "What is SCIM" is a paragraph the model already knows.
Does anyone ask for you by name?Yes. Branded and "site:reddit" demand is enormous.Rarely. Branded volume is a rounding error against category terms.
Is there a contract and a renewal date?Yes. Reportedly $60M a year.No. There is nothing to negotiate and nobody to negotiate with.

Take those one at a time. If your best page explains what SCIM is, you are competing with a paragraph the model can generate from memory. Branded query volume is the cleanest proxy for whether your audience is yours or Google's, and if nobody types your name, the crawler is not the one who owns the relationship. And the third question is the one that actually decides it: what breaks in the first ninety days? Publishers ask this and get an answer in ad revenue. In B2B SaaS the answer is usually pipeline that started with an unbranded search eleven months before a contract was signed, which almost nobody can attribute properly.

Blocking works as a strategy when scarcity is on your side. It works as theater the rest of the time.

B2B SaaS was never in the traffic business

This is where the publisher story stops being our story.

A publisher converts an impression into an ad. Volume is the product. Take away half the visitors and you have taken away half the revenue, mechanically.

We convert a specific person with a specific problem into a purchase that happens weeks or quarters later. We never needed volume. We needed the right five hundred people, and we bought their attention with traffic because traffic was the cheapest way to reach them. The traffic was the delivery mechanism, not the value.

Which means the zero-click web is a smaller problem for us than for a newsroom, and the citation is worth more than the click. When an answer engine names you inside a comparison of identity vendors, you are in the consideration set before the buyer has a session, a cookie or a form fill. That is the job the analyst grid used to do, and companies paid a great deal for it.

The mistake is to keep measuring the delivery mechanism after the value moved. Sessions are a lagging, increasingly meaningless number. I wrote the complete guide to generative engine optimization because I kept meeting teams who were winning citations and reporting a traffic decline to their board as a failure.

What I am doing instead

Six things, in the order I would do them.

1. Audit the crawl before touching a setting. Pull thirty days of bot logs and thirty days of assistant referrals. You want the crawl-to-referral ratio per bot. Most B2B sites I have looked at are still under 1% of sessions from AI assistants, which means switching training crawlers to charge or block costs almost nothing in reach and starts a paper trail that matters if the licensing market matures. Do this before September 15, because if you are on Cloudflare a default may make the decision for you.

2. Build the things a model cannot summarize away. My hashing tools turn impressions into clicks at a multiple of what my article pages manage, and they are most of the reason that site-wide 3.14% is not a fraction of a percent. Same author, same topics, different asset class. A model can compress my comparison of Argon2, bcrypt, scrypt and PBKDF2 into four sentences. It cannot hash your string for you. Interactive assets are the closest thing to a moat that content people have left.

3. Publish primary material. Original data, methodology you show your work on, dated first-person accounts of things you actually ran. The correction fiasco this week is the proof: the sites that got it wrong were the ones with nothing of their own to add. My GEO market research and the 50,000-citation measurement study get cited far more often than anything I have written that explains a concept, because there is nowhere else to get them.

4. Structure for extraction. Answer first, one claim per paragraph, verifiable numbers with sources, schema markup, an llms.txt file, dated methodology pages. None of this is exotic. Most sites still do none of it. The AEO strategy playbook has the full checklist, and the scoring model behind it lives in GEO Compass.

5. Own a distribution channel nobody can deprecate. Email list, community, podcast, the boring durable things. Every publisher in that Journal story that is doing well is doing this. People Inc. now gets a quarter of its traffic from Google, down from more than half two years ago, and grew revenue anyway through events, social and apps.

6. Change the metric you report. Citation share by engine, prompt coverage for the questions your buyers actually ask, branded search volume. I track how each engine represents me, including the ones people underestimate, which is why I keep a running breakdown of how Grok retrieves and cites. Sessions go at the bottom of the deck now, if at all. I stopped paying for a rank tracker over this, and wrote down why, which is a slightly awkward thing to admit in a post whose central numbers come from Semrush.

The uncomfortable part

If the large publishers actually follow through, the open web that models ground their answers on gets thinner and stranger. What remains will skew toward whoever kept the door open: platforms with licensing contracts, SEO farms, LinkedIn, and independent practitioners publishing first-hand material.

That last group is a real opening, and I am part of it, so I want to be honest that I benefit from a situation I do not think is healthy. An information commons where the highest-quality sources have withdrawn behind registration walls and the models are grounding on whatever is left is not a better web. It is just a differently broken one. Nieman Lab's read of the same week is more sympathetic to the publishers than mine, and worth sitting with, because the people making these calls are not being irrational. They are being cornered.

Time's approach, writing text-based ads that human readers never see but crawlers might pick up, tells you where this is heading. Their COO described it as serving two audiences now. He is right, and I do not love what that implies about the second one.

Here is the question I would ask your team this week, and I mean it as a real question rather than a closing line. If Google sent you zero clicks next quarter but cited you in every answer in your category, would your pipeline be better or worse?

Most teams cannot answer that. It took me a year of instrumentation to answer it for myself, and the answer surprised me.

FAQ

Are publishers really about to block Google?

Some are weighing it. Reddit has internally discussed ending Google's access for AI use as its reported $60 million a year deal comes up for renewal, and USA Today Co., Politico, Reuters, the Economist and People Inc. are all reassessing the relationship. None has blocked Googlebot outright. The threat is currently negotiating pressure ahead of contract renewals and regulatory changes.

How much has Google search traffic to publishers actually fallen?

The Wall Street Journal corrected its figures on July 22, 2026. Comparing U.S. organic Google search traffic in June 2025 to June 2026, USA Today's national paper was down 18%, Politico 20%, and CNN and Business Insider 31% each, per Semrush. Trailing twelve-month averages showed declines from 3.2% to 44% across major titles, while the Guardian and BBC both grew.

Can a site block Google's AI features without losing search rankings?

Not reliably, yet. Google-Extended keeps content out of Gemini training but does not stop Googlebot from crawling for AI Overviews and AI Mode. Nosnippet removes preview text from ordinary results too. The U.K. CMA ordered Google on June 3, 2026 to give publishers an AI-features opt-out that does not affect search rankings, and Google said it will comply, starting with a U.K. test and no announced global date.

What changes on September 15, 2026?

Cloudflare's new defaults block mixed-use crawlers on pages carrying ads for new customers, new sites on existing accounts and free-tier accounts. Crawlers that combine search, agent and training behavior are treated under the most restrictive rule the site owner has set, which includes Googlebot. Site owners can opt out in the dashboard before that date.

Should a B2B SaaS company block AI crawlers?

Usually not the search crawler, and rarely all of them. Blocking makes sense when your content cannot be substituted and you have a contract to negotiate. For most B2B SaaS sites the better sequence is to audit crawl-to-referral ratios per bot, set training crawlers to charge or block if assistant referrals are negligible, keep search open, and shift measurement from sessions to citation share.

What should a B2B SaaS team do before September 15, 2026?

Open the bot settings on your CDN and decide deliberately rather than inheriting a default. Pull per-bot crawl logs and assistant referral counts, set training and agent crawlers to charge or block if referrals are negligible, keep the search crawler allowed, and add citation share to the reporting deck so the board sees the metric that now carries the pipeline.

Get the newsletter

New writing on identity, AI security, and building software, delivered when it ships. No tracking pixels, no funnels, unsubscribe with one click.