Robots.txt has answered one question for thirty years: may you fetch this file. It was never designed to answer the question that matters now, which is what you may do with the file once you have it. Cloudflare's Content Signals Policy is the first serious attempt to bolt that second question onto the first, and the way it shipped tells you more than the specification does.
The mechanism is deliberately unglamorous. Cloudflare added a line to robots.txt carrying up to three named signals, each set to yes or no, alongside a plain-English policy block explaining what the signals mean. The syntax in Cloudflare's own developer documentation reads `Content-signal: search=yes, ai-train=no, use=reference`. No new file, no new standard body, no crawler changes required.
What makes it worth an hour of an enterprise SEO team's attention is not the syntax. It is that Cloudflare set the defaults on behalf of millions of sites, and that the field it chose to leave empty is the one every generative engine optimization program is built around.
Four numbers, one argument. A large automatic rollout, a three-field specification, a default that fills in two fields, and roughly a year of operating history. The gap between the second and third numbers is the whole piece.
What the Content Signals Policy actually adds to robots.txt
The Content Signals Policy separates permission to crawl from permission to use, which is a distinction robots.txt has never carried. A crawler reads the same directives it always did, then reads a second line describing what the operator may legitimately do with what it fetched.
That separation matters because the crawler population stopped being homogeneous. A single bot can fetch a page to build a search index, to ground a live generative answer, or to accumulate training corpus, and those are three different commercial relationships with the publisher. Blocking by user agent, the approach most sites still run, forces a single yes or no across all three. We walked through why that collapses in practice in the difference between training and retrieval crawlers, and the Content Signals Policy is the first widely deployed attempt to give the three uses separate switches.
| SIGNAL | WHAT IT GOVERNS, PER CLOUDFLARE'S DOCUMENTATION | WHAT IT MAPS TO COMMERCIALLY | MANAGED DEFAULT |
|---|---|---|---|
| search | Building a search index and returning hyperlinks and short excerpts from the site | The classic bargain: you index me, you send me traffic | yes |
| ai-input | Feeding content into AI models, including retrieval augmented generation, grounding, and real-time use in generative answers | AI Overviews, AI Mode, Perplexity, ChatGPT search. Citation with little or no click | unset |
| ai-train | Training or fine-tuning AI models | Corpus acquisition. No attribution, no traffic, no ongoing relationship | no |
Read the middle column and the commercial logic of the default becomes obvious. Search is a trade publishers have accepted for decades. Training is a transfer with nothing flowing back. Those two are easy calls, and Cloudflare made them. The middle row is the one where reasonable publishers genuinely disagree, so Cloudflare declined to decide it for them.
The default reached 3.8 million domains with ai-input left blank
Cloudflare applied the policy automatically to every domain using its managed robots.txt feature, a population the company put at more than 3.8 million sites, and wrote search=yes and ai-train=no into each of them without the site owner taking any action.
Defensible on the merits, and still worth noticing as a governance event. A vendor made a rights declaration on behalf of millions of publishers, in a file those publishers rarely open, expressing a position most of them had never been asked about. It is the same shape as every other default set at the platform layer: the vendor integrates once, on behalf of everybody, and the position propagates to people who never evaluated it.
None of this is a criticism of the design. A policy that shipped only to sites that opted in would have reached a rounding error of the web and proved nothing. Shipping it on by default is what made it a real-world artifact rather than a proposal. The cost of that choice is that a lot of publishers now carry a stated position they did not author.
Google says the Content Signals Policy changes nothing for Google
Google responded to the Content Signals Policy by saying the directive has no effect on Google whatsoever, a position Search Engine Roundtable reported and Google has not softened since. Googlebot reads robots.txt the way it always has and ignores the content signals line entirely.
The reasoning is consistent with Google's long-held position on unofficial robots.txt extensions, which is the same reasoning it gives for why llms.txt does nothing for Google Search. Google honors the official standard and the mechanisms it publishes itself, notably the Google-Extended token for training and the standard disallow rules for crawling. It does not honor directives invented by third parties, however sensible.
“A signal that the largest consumer of your content ignores is not worthless. It is worthless against that one consumer, which is a much narrower claim than the one most people take away.”
Take the narrow claim seriously, because it cuts both ways. Google ignoring the line means you cannot use ai-input=no to keep yourself out of AI Overviews; the mechanism for that remains Google's own controls, and using them still costs you the classic snippet. It also means you cannot use ai-input=yes to buy your way into AI Mode. Google's inclusion decisions are unaffected in both directions.
What the line does reach is everyone else. Perplexity, OpenAI's search crawler, Anthropic's search crawler and the long tail of retrieval operators publish compliance postures that treat robots.txt as authoritative, and several have signaled they will read usage signals where present. Those operators are a minority of your referred traffic today and a growing share of the answers your buyers read. The realistic assessment is that content signals are advisory everywhere, ignored by the largest player, and read by a meaningful set of the rest.
An unset ai-input signal is the worst of the three available states
A publisher choosing among ai-input=yes, ai-input=no and leaving it unset faces three options where the first two are coherent strategies and the third is an absence of one. Ambiguity in a machine-readable file resolves in favor of whoever is reading it.
The case for yes is straightforward for most commercial sites. You want to be the source an engine names when a buyer asks it to shortlist vendors in your category. That is the entire premise of a generative engine optimization program, and declaring the use you are actively courting costs you nothing.
The case for no is real for a narrower set: publishers whose revenue is the click, subscription businesses whose archive is the product, and rights holders assembling a record for licensing negotiations or litigation. Cloudflare's documentation explicitly frames the signals as express reservations of rights under Article 4 of European Union Directive 2019/790, which is a legal posture, not a traffic tactic. If that is why you are here, a documented, dated, machine-readable reservation has value regardless of whether any given crawler honors it.
The case for leaving it blank is that you have not had the conversation. That is the state 3.8 million domains are in, and it produces the worst outcome of the three: no traffic protection, no legal record, and no declared eligibility, while your team believes Cloudflare handled it. The same pattern shows up whenever a robots directive is inherited rather than chosen, which is how sites end up blocking AI crawlers by accident.
How to set your content signals deliberately
Setting content signals deliberately takes about an hour and touches three systems: the Cloudflare dashboard that generates the file, the robots.txt itself, and whatever internal document records why the position was chosen.
Then measure the thing the signals are meant to affect, rather than the signals themselves. Robots directives are inputs, and the only output that matters is whether you appear in generative answers. The generative AI reporting Google now publishes covers Google's surfaces, which are precisely the surfaces content signals do not reach, so pair it with citation monitoring on the engines that do read the line. If a change to ai-input moves anything, it will move there first.
Questions teams are asking about the Content Signals Policy
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.