Something Inc.Schedule a free consultation
TECHNICAL SEO

Block Googlebot or Not? Cloudflare's Default Kicks In Sept 15

Starting September 15, Cloudflare will block Googlebot by default for any new domain that blocks AI training crawlers — because Google won't split the two. Here's the actual decision, not the panic version.

JBJosh BernsteinManaging Partner · AUG 4, 2026 · 8 MIN READ

Cloudflare published its policy on July 1, 2026, and a Hacker News thread surfaced it again this morning, Aug 4, with 51 points and 28 comments arguing over what it means. The headline version going around is wrong. Cloudflare is not blocking Google. Cloudflare is changing what happens by default when a new site asks it to block AI training crawlers, and Googlebot happens to be one of them, whether you meant to block it or not.

KEY TAKEAWAYSeptember 15 is not a decision Cloudflare is making for you. It's the date a default flips for any new domain. If you block AI training crawlers on a new Cloudflare site and don't check one setting, Googlebot goes down with GPTBot. Most site owners have never heard of this and will find out from a traffic drop.

What Cloudflare's block Googlebot default actually changes

Cloudflare's policy, published under the banner "Content Independence Day" on July 1, 2026, sorts every bot hitting a site into three buckets: Search, Agent, and Training. Starting September 15, on pages a site monetizes with ads, the new defaults are straightforward on paper. Training bots are blocked. Agent bots are blocked. Search bots are allowed. Cloudflare's own reasoning, from the announcement: "on those pages, we treat human attention as the end goal, and keep away the bots that may prevent this attention." That is a reasonable position for a publisher worried about AI models training on its content for free while sending it fewer clicks in return.

The complication is the middle category most people skip past. A handful of crawlers don't fit neatly into one bucket, because the same company runs them for more than one purpose. Cloudflare names three explicitly: Googlebot, Applebot, and BingBot. These are multi-purpose crawlers, and Cloudflare's policy applies what it calls "the most restrictive applicable rules" to them, based on everything they do, not just the one function you had in mind when you flipped a switch.

Why blocking Googlebot happens without anyone deciding to

Here's the mechanism. Google uses Googlebot both to crawl for classic search indexing and, per Cloudflare's classification, for AI training purposes tied to Gemini. It is not two separate bots with two separate user agents you can selectively allow. It is one crawler serving two functions. So when a new domain's owner blocks Training bots, intending to stop their content from feeding a model they never agreed to train, the most-restrictive rule kicks in and Googlebot gets blocked outright, taking classic Google Search indexing down with it. Cloudflare states this plainly in its own documentation: "multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training."

"Multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training." — Cloudflare, Content Independence Day policy, July 1, 2026.

The Hacker News thread discussing this today split roughly where you'd expect. One commenter, simonw, framed it as Google forcing a binary choice on publishers: keep Googlebot's search crawl and accept Gemini training exposure, or protect your content from training and lose Google Search visibility as a side effect, because Google refuses to offer a crawler that does one without the other. Other commenters read it the opposite way, as leverage finally shifting toward publishers after years of one-sided scraping, on the theory that enough sites making this exact tradeoff visibly is what eventually forces Google to split the two functions. Both readings can be true at once. The business logic point stands regardless of which side you're rooting for: Google chose not to separate the two functions, and that choice, not Cloudflare's policy, is what puts your search visibility on the table.

Who this actually affects

The scope is narrower than the alarmed version of this story suggests, and the details matter more than the headline. It's also worth remembering this isn't the first time a crawler-access change has caught sites by surprise: our own coverage of sites blocking AI crawlers by accident found 30% of sites were already unintentionally blocking GPTBot through stale rules nobody had reviewed in years, well before Cloudflare's new default made the same category of accident possible for Googlebot specifically. First, the new default only applies to domains newly onboarding to Cloudflare after September 15. If your site is already on Cloudflare today, this default does not retroactively flip your existing settings; you're grandfathered into whatever you already have configured. Second, the restrictive default only bites on pages you've marked as ad-monetized. A gated SaaS product with no display advertising, for instance, sits outside the specific rule discussed here. Third, and most important for anyone launching a new property on Cloudflare after the cutover: the block only triggers if you choose to block Training bots. Leave Training allowed, and Googlebot is unaffected by this specific rule, though you've then accepted the AI-training-scrape tradeoff you may have been trying to avoid in the first place.

IF A NEW DOMAIN SETS…GOOGLEBOT (SEARCH)GPTBOT / CLAUDEBOT (TRAINING)NET EFFECT
Block Training only (the trap)Blocked, as a side effectBlockedGoogle Search visibility lost without ever choosing to lose it
Block Training + explicitly allow Search exceptionAllowedBlockedAI training access denied, classic search crawl preserved
Allow everything (do nothing)AllowedAllowedNo protection from AI training scrape, no search-visibility risk
Block Training + Agent + SearchBlockedBlockedFull lockout — defensible only for gated, non-SEO-dependent properties

The crawler-economics case for treating Googlebot differently from GPTBot or ClaudeBot in the first place is not new, and it's worth restating here because it's exactly the logic this new default silently overrides for anyone who isn't paying attention. Our own technical audits work has tracked crawl-to-referral ratios across bots for months: Googlebot crawls a page and sends back meaningful referral traffic at a ratio close to 5:1, while pure AI-training crawlers run orders of magnitude worse. The two are not the same cost-benefit trade, and a single "block AI crawlers" switch that can't tell them apart is exactly the kind of blunt instrument that creates accidents like this one.

~5:1
Googlebot's crawl-to-referral ratio — requests sent per referral it sends back
~5,143:1
ClaudeBot's ratio, per TechnologyChecker.io/Cloudflare Radar data
~38,000:1
Anthropic's training-crawl ratio, per Cloudflare/ppc.land data
READING THE RATIOSThose figures were previously reported in our own crawler-economics coverage. Googlebot sits nowhere near the crawlers an AI-training-blocking policy is meant to target. It is not the crawler anyone worried about uncompensated AI scraping should be losing sleep over — which is exactly why an accidental block is such an expensive mistake to make by default.

The decision: block Googlebot or carve out an exception

1If you're launching a new domain on Cloudflare after Sept 15Before you block Training bots, go into Security settings and explicitly confirm you want "no changes on Training crawlers that also crawl for Search purposes" — Cloudflare's own opt-out language for this exact scenario. Do this the same day you configure the domain, not after your search traffic drops and you go looking for why.
2If you're already live on an existing Cloudflare domainYou're grandfathered in for now, but confirm your current bot-management settings actually say what you think they say. Don't assume; check. Defaults change, and a policy you set eighteen months ago under a different UI may not map cleanly to today's Search/Agent/Training categories.
3If you're deciding whether to block AI training crawlers at allSeparate the decision from the accident. Blocking Training crawlers to protect your content from uncompensated AI training is a legitimate, defensible call, especially with Cloudflare's pay-per-citation infrastructure now giving you a monetization lever instead of a binary block. Just make that decision on purpose, with Googlebot carved out explicitly, rather than by default.
4If you manage a portfolio of sites across agencies or business unitsAudit every property individually rather than pushing one bot-management policy fleet-wide. Ad-monetized content sites, gated SaaS products, and lead-gen landing pages have different real exposure to this rule, and a single blanket policy is how the accident spreads across a portfolio instead of staying contained to one property. This is exactly the kind of technical-access review we build into new site builds from day one, rather than bolting on after a launch.

None of this is a reason to panic about Cloudflare, or to conclude that blocking AI crawlers is a bad idea. It's a reason to read the actual mechanism instead of the headline. The debate over whether to block AI crawlers or charge for their traffic is a real, ongoing one, and Cloudflare has kept iterating on both sides of it. This specific policy detail, though, isn't really about that debate. It's a plumbing detail with a hard deadline, and plumbing details are exactly the kind of thing that get missed by teams focused on the bigger strategic question and not the settings page underneath it.

It's also a preview of a pattern that's going to keep recurring. As crawler-management platforms build increasingly granular controls for an increasingly granular set of AI-related traffic types, the number of settings a site owner is implicitly making decisions about, whether they know it or not, keeps climbing. Bot management used to mean one blunt toggle: block or allow. Now it means a taxonomy of purposes layered on top of a taxonomy of vendors, evaluated against a taxonomy of monetization states per page. Every layer of that complexity is a new place for a default to quietly make a decision on your behalf. The specific fix here is a two-minute settings check. The durable fix is treating every new crawler-management rollout, from any vendor, as something worth reading the actual documentation on before the deadline, not after the traffic graph moves.

DO THIS NEXTIf you're launching anything new on Cloudflare between now and September 15, open Security → Bots today, not the week your organic traffic dips. Confirm the Training/Search exception explicitly. It's a two-minute setting. The alternative is diagnosing a Google Search Console traffic drop three weeks from now and tracing it back to a checkbox nobody knew existed.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

JB
Josh BernsteinMANAGING PARTNER, SOMETHING INC.

Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.