Somewhere on your team, someone has probably already argued for blocking AI crawlers outright. It feels like the obvious defensive move: AI companies train on your content and send almost nothing back, so cut them off at robots.txt and stop the bleeding. A Wharton and Rutgers study published this year says that instinct is not just wrong, it's actively costing the publishers who act on it real, measured traffic, for zero protection in return.
The instinct to block AI crawlers, and why it's understandable
The case for blocking AI crawlers sounds reasonable on its face. GPTBot, ClaudeBot, and their peers request your content, most of it gets folded into training runs or live retrieval you don't control, and the traffic you get back in exchange is thin at best. Server logs across the industry show AI and LLM indexing bots quadrupling their share of traffic in eight months, from 2.6% to 10.1%, and some crawlers running crawl-to-refer ratios above 700-to-1, meaning hundreds of requests for every visit they send back. If you're paying the bandwidth bill and getting almost nothing in return, blocking looks like the rational move.
That's the reasoning behind a lot of the robots.txt disallow rules that went up across the publishing industry over the past year. It's also, according to the best available data, backwards.
It's worth being fair to the instinct before dismantling it. Server-side, the complaint isn't imaginary: sites on serverless infrastructure can pay $30 to $300 a month just serving bot requests once volume climbs into six figures a day, and larger properties report 40-60% of total traffic now coming from bots rather than humans. That's a real line item, and it's reasonable for an ops team to want it gone. The mistake isn't noticing the cost. It's assuming a robots.txt block is the tool that removes it, when the actual research on what happens after a block goes up says something different.
What the Wharton and Rutgers data actually shows
"Strategic Response of News Publishers to Generative AI," authored by Hangcheng Zhao of Rutgers Business School and Ron Berman of The Wharton School, is the most rigorous look yet at what actually happens to a publisher's traffic after it blocks AI crawlers. The paper's own numbers moved between versions, which is itself informative: the earlier December 2025 draft, focused on large publishers, found a 23.1% drop in monthly visits after blocking. The April 2026 revision, with a broader and more current sample, settled on a 7% weekly traffic loss within six weeks of a block going into effect.
Read the two figures side by side and the direction, not just the magnitude, is the finding that matters. Whether the real number is 7% or 23%, both revisions of the same research point the same way: publishers that blocked AI crawlers lost measurable traffic, and neither version of the study found an offsetting gain anywhere else in the business to justify it.
The revision between drafts is also a useful lesson in reading research honestly instead of cherry-picking the scarier number. The 23.1% figure circulated first and still shows up in older coverage, but it came from an early sample weighted toward large publishers specifically, the outlets with the most brand recognition to lose and the most alternate paths for readers to find them anyway. The later, broader 7% figure is the more defensible number to plan around, and it's still not zero. A real, repeatable traffic loss with no offsetting benefit is a bad trade regardless of which version of the paper you're citing.
Blocking doesn't even stop the citation
The part of this that should sting most for anyone who blocked crawlers specifically to protect their content is that it usually doesn't work, even on its own narrow terms. BuzzStream research published March 19, 2026 found that 70.6% of news sites that had declared a block were still getting cited by AI engines anyway.
This lines up with what we found when we looked at whether robots.txt actually stops AI crawlers: a July 2026 study of over 10,000 domains found 39.5% of declared GPTBot bans went unhonored outright, and Cloudflare has publicly accused Perplexity of spoofing its user-agent specifically to keep scraping sites that blocked it. Between crawlers that ignore the file and citations sourced from anywhere else your facts appear, a block is a much leakier barrier than it feels like when you're writing the Disallow line.
| WHAT A BLOCK IS MEANT TO DO | WHAT THE DATA SHOWS |
|---|---|
| Stop the content from being used in AI answers | 70.6% of blocking sites got cited anyway, per BuzzStream's March 2026 audit |
| Protect referral traffic by forcing a fair exchange | No offsetting traffic gain found in either Zhao/Berman study revision |
| Cost nothing since the crawler wasn't sending much traffic anyway | Publishers still lost 7% of weekly traffic within six weeks of blocking |
Why the traffic loss happens even without a single scrape
The mechanism Zhao and Berman point to isn't some indirect algorithmic penalty. It's simpler and more structural than that. When a publisher blocks a crawler, its material stops surfacing in that engine's AI-generated summaries and answer tools. A brand that isn't named in the answer is a brand fewer people encounter at all, not just a brand that loses one referral click. Readers who would have discovered the outlet through an AI-generated summary, then clicked through later for depth, through search, through a bookmark, through recognizing the name next time, never get that first exposure in the first place.
That's a brand-awareness loss wearing a technical-SEO costume, and it compounds the same way comparison content compounds: being present where buyers are asking questions builds the recognition that makes every other channel work better. Blocking the crawler doesn't defend that recognition. It quietly switches it off.
There's also a competitive angle the Zhao and Berman paper doesn't need to spell out directly, because it follows from the same mechanism. If your competitor's content stays reachable while yours doesn't, the engine doesn't leave a gap where your citation used to be. It fills that gap with whoever's still answering the question, which in most categories means a direct competitor. Blocking a crawler unilaterally, in a category where competitors haven't done the same, doesn't remove you from the conversation. It hands your seat in that conversation to someone else, for free, while you're still paying to serve the bot traffic you didn't manage to keep out anyway.
What to do instead of a blanket block
None of this means AI crawler traffic is free to ignore, or that every crawler deserves unrestricted access. It means the blanket block, the reflexive "disallow everything with 'GPT' or 'AI' in the name" approach, is the wrong tool for the actual problem.
If you're running a technical SEO program and the instinct on your team is to block first and ask questions later, this is the study to bring to that conversation. A technical access audit that checks what's actually happening in your server logs, which crawlers are honoring your rules, which are getting cited through other paths regardless, and where real server cost is coming from, will tell you more than a blanket policy ever will. For devtools and other technical categories where AI-assisted research is already the default entry point, the Wharton and Rutgers numbers aren't an edge case. They're the baseline cost of getting this decision wrong, and right now the data says most publishers are getting it wrong in the direction that costs them the most.
The framing worth carrying out of this study isn't "never block anything." It's that a block is a cost with a specific, now-measured price tag attached, not a free defensive move with no downside. Once you know the price, a blanket block against every AI crawler stops making sense on its own terms, and the decision gets a lot more specific: which crawler, on which pages, for what actual reason. A pricing page you genuinely don't want scraped and summarized is a different call than your entire editorial archive, and treating them exactly the same way is how a publisher ends up eating the full 7% weekly traffic cost for a protection benefit that BuzzStream's own audit says arrives less than a third of the time.
None of this is an argument that AI crawlers deserve unconditional trust either. It's an argument for measuring before you act, the same discipline that governs every other technical SEO decision that touches revenue. Pull your own logs before writing the Disallow line, check which crawlers are honoring your existing rules and which ones aren't, and decide per-crawler and per-section rather than reaching for the blanket rule that feels satisfying in the moment and, per the best data available right now, costs more than it protects.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.