Something Inc.LoginSchedule a free consultation
TECHNICAL SEO

The Content Signals Policy Leaves Your Best Signal Blank

Cloudflare's Content Signals Policy writes three AI usage directives into robots.txt and sets two of them for you. The one it leaves unset, ai-input, is the only one that governs whether your pages can appear inside an AI answer.

TECHNICAL SEOAI CRAWLERSSEP 2026

Robots.txt has answered one question for thirty years: may you fetch this file. It was never designed to answer the question that matters now, which is what you may do with the file once you have it. Cloudflare's Content Signals Policy is the first serious attempt to bolt that second question onto the first, and the way it shipped tells you more than the specification does.

The mechanism is deliberately unglamorous. Cloudflare added a line to robots.txt carrying up to three named signals, each set to yes or no, alongside a plain-English policy block explaining what the signals mean. The syntax in Cloudflare's own developer documentation reads `Content-signal: search=yes, ai-train=no, use=reference`. No new file, no new standard body, no crawler changes required.

What makes it worth an hour of an enterprise SEO team's attention is not the syntax. It is that Cloudflare set the defaults on behalf of millions of sites, and that the field it chose to leave empty is the one every generative engine optimization program is built around.

3.8M+
domains running Cloudflare's managed robots.txt, where Cloudflare inserted the policy and set the signals on the site owner's behalf, per Cloudflare's launch announcement
3
signals defined by the policy: search, ai-input and ai-train, each expressing a permitted use rather than a permitted fetch
2
of those three set by the managed default. search=yes and ai-train=no were applied; ai-input was deliberately left neutral
SEP 24
the 2025 date Cloudflare announced the Content Signals Policy. Its developer documentation was last revised on August 3, 2026

Four numbers, one argument. A large automatic rollout, a three-field specification, a default that fills in two fields, and roughly a year of operating history. The gap between the second and third numbers is the whole piece.

What the Content Signals Policy actually adds to robots.txt

The Content Signals Policy separates permission to crawl from permission to use, which is a distinction robots.txt has never carried. A crawler reads the same directives it always did, then reads a second line describing what the operator may legitimately do with what it fetched.

That separation matters because the crawler population stopped being homogeneous. A single bot can fetch a page to build a search index, to ground a live generative answer, or to accumulate training corpus, and those are three different commercial relationships with the publisher. Blocking by user agent, the approach most sites still run, forces a single yes or no across all three. We walked through why that collapses in practice in the difference between training and retrieval crawlers, and the Content Signals Policy is the first widely deployed attempt to give the three uses separate switches.

SIGNALWHAT IT GOVERNS, PER CLOUDFLARE'S DOCUMENTATIONWHAT IT MAPS TO COMMERCIALLYMANAGED DEFAULT
searchBuilding a search index and returning hyperlinks and short excerpts from the siteThe classic bargain: you index me, you send me trafficyes
ai-inputFeeding content into AI models, including retrieval augmented generation, grounding, and real-time use in generative answersAI Overviews, AI Mode, Perplexity, ChatGPT search. Citation with little or no clickunset
ai-trainTraining or fine-tuning AI modelsCorpus acquisition. No attribution, no traffic, no ongoing relationshipno

Read the middle column and the commercial logic of the default becomes obvious. Search is a trade publishers have accepted for decades. Training is a transfer with nothing flowing back. Those two are easy calls, and Cloudflare made them. The middle row is the one where reasonable publishers genuinely disagree, so Cloudflare declined to decide it for them.

THE DISTINCTION WORTH HOLDINGThis is a preference expression, not an access control. Cloudflare's documentation is explicit that robots.txt compliance is voluntary and that some operators will crawl regardless. If your goal is to stop a bot, content signals are the wrong tool and a WAF rule is the right one. If your goal is to state a position the compliant operators can read, and to create a record, this is exactly the tool.

The default reached 3.8 million domains with ai-input left blank

Cloudflare applied the policy automatically to every domain using its managed robots.txt feature, a population the company put at more than 3.8 million sites, and wrote search=yes and ai-train=no into each of them without the site owner taking any action.

Defensible on the merits, and still worth noticing as a governance event. A vendor made a rights declaration on behalf of millions of publishers, in a file those publishers rarely open, expressing a position most of them had never been asked about. It is the same shape as every other default set at the platform layer: the vendor integrates once, on behalf of everybody, and the position propagates to people who never evaluated it.

01The easy calls got made, the hard one did notsearch=yes and ai-train=no are the two positions a large majority of publishers would pick unprompted. Leaving ai-input neutral is Cloudflare correctly declining to guess on the only genuinely contested question.
02Neutral is not the same as noAn unset signal does not deny anything. It declines to state a preference, which leaves the operator to infer one. Publishers reading a headline about Cloudflare blocking AI crawlers frequently believe they are covered for generative answers. They are not.
03The marketing team and the infrastructure team want opposite things hereai-train=no protects the asset. ai-input=yes is what makes the asset eligible to be cited in an answer. Those sit in one line of one file, and in most organizations they are owned by two departments that have never discussed it.
04A default applied once keeps applyingManaged robots.txt regenerates. A team that hand-edits the file without moving the setting upstream in the Cloudflare dashboard will find their change reverted on the next regeneration, which is the ordinary failure mode for every managed configuration file.

None of this is a criticism of the design. A policy that shipped only to sites that opted in would have reached a rounding error of the web and proved nothing. Shipping it on by default is what made it a real-world artifact rather than a proposal. The cost of that choice is that a lot of publishers now carry a stated position they did not author.

Google says the Content Signals Policy changes nothing for Google

Google responded to the Content Signals Policy by saying the directive has no effect on Google whatsoever, a position Search Engine Roundtable reported and Google has not softened since. Googlebot reads robots.txt the way it always has and ignores the content signals line entirely.

The reasoning is consistent with Google's long-held position on unofficial robots.txt extensions, which is the same reasoning it gives for why llms.txt does nothing for Google Search. Google honors the official standard and the mechanisms it publishes itself, notably the Google-Extended token for training and the standard disallow rules for crawling. It does not honor directives invented by third parties, however sensible.

“A signal that the largest consumer of your content ignores is not worthless. It is worthless against that one consumer, which is a much narrower claim than the one most people take away.”

Take the narrow claim seriously, because it cuts both ways. Google ignoring the line means you cannot use ai-input=no to keep yourself out of AI Overviews; the mechanism for that remains Google's own controls, and using them still costs you the classic snippet. It also means you cannot use ai-input=yes to buy your way into AI Mode. Google's inclusion decisions are unaffected in both directions.

What the line does reach is everyone else. Perplexity, OpenAI's search crawler, Anthropic's search crawler and the long tail of retrieval operators publish compliance postures that treat robots.txt as authoritative, and several have signaled they will read usage signals where present. Those operators are a minority of your referred traffic today and a growing share of the answers your buyers read. The realistic assessment is that content signals are advisory everywhere, ignored by the largest player, and read by a meaningful set of the rest.

An unset ai-input signal is the worst of the three available states

A publisher choosing among ai-input=yes, ai-input=no and leaving it unset faces three options where the first two are coherent strategies and the third is an absence of one. Ambiguity in a machine-readable file resolves in favor of whoever is reading it.

The case for yes is straightforward for most commercial sites. You want to be the source an engine names when a buyer asks it to shortlist vendors in your category. That is the entire premise of a generative engine optimization program, and declaring the use you are actively courting costs you nothing.

The case for no is real for a narrower set: publishers whose revenue is the click, subscription businesses whose archive is the product, and rights holders assembling a record for licensing negotiations or litigation. Cloudflare's documentation explicitly frames the signals as express reservations of rights under Article 4 of European Union Directive 2019/790, which is a legal posture, not a traffic tactic. If that is why you are here, a documented, dated, machine-readable reservation has value regardless of whether any given crawler honors it.

The case for leaving it blank is that you have not had the conversation. That is the state 3.8 million domains are in, and it produces the worst outcome of the three: no traffic protection, no legal record, and no declared eligibility, while your team believes Cloudflare handled it. The same pattern shows up whenever a robots directive is inherited rather than chosen, which is how sites end up blocking AI crawlers by accident.

THE TEST THAT SETTLES ITAsk whoever owns the revenue number one question: if ChatGPT names three vendors in our category next week, do we want to be one of them? If the answer is yes, ai-input=yes is the only setting consistent with the rest of the budget. If the answer is that the click is the product and citation without a visit is a loss, set it to no and mean it. Either answer is defensible. Not answering is not.

How to set your content signals deliberately

Setting content signals deliberately takes about an hour and touches three systems: the Cloudflare dashboard that generates the file, the robots.txt itself, and whatever internal document records why the position was chosen.

Read your own robots.txt before the meetingFetch yourdomain.com/robots.txt and look for the Content-signal line and the policy comment block. Most teams discover the policy is already present and already says something. Establish what you currently declare before debating what you should declare.
Decide ai-input on revenue model, not on sentimentClick-monetized publishing and subscription archives have a genuine case for no. Lead-generation, SaaS, services and considered-purchase ecommerce almost always want yes, because being named in an answer is the acquisition event. Sentiment about AI companies is not an input to this.
Change it upstream, not in the fileIf Cloudflare manages your robots.txt, edit the setting in the dashboard. A hand-edit of a generated file survives until the next regeneration and then quietly disappears, taking your stated position with it.
Keep ai-train separate from ai-inputThese are different transactions and deserve different answers. ai-train=no with ai-input=yes is the coherent position for most commercial sites: do not absorb my archive into a model, do cite me when you answer a question about my category.
Confirm your access rules agree with your signalsA signal that says yes while a firewall rule blocks the same crawler is a contradiction the crawler resolves by leaving. Reconcile the two, which is standard scope in an SEO and GEO audit and the single most common finding when citation volume has quietly fallen.

Then measure the thing the signals are meant to affect, rather than the signals themselves. Robots directives are inputs, and the only output that matters is whether you appear in generative answers. The generative AI reporting Google now publishes covers Google's surfaces, which are precisely the surfaces content signals do not reach, so pair it with citation monitoring on the engines that do read the line. If a change to ai-input moves anything, it will move there first.

DO THIS NEXTPull your robots.txt today and write down what the Content-signal line currently says, including whether ai-input is present at all. Take that one line to whoever owns pipeline and get an explicit yes or no on it inside a week. Cloudflare's content signals documentation is short enough to read in the meeting. The setting is trivial to change and the position it states is not, which is the usual reason it never gets changed.

Questions teams are asking about the Content Signals Policy

Is the Content Signals Policy already on my site?If your robots.txt is managed by Cloudflare, very likely yes. Cloudflare applied it to more than 3.8 million domains automatically. Fetch yourdomain.com/robots.txt and look for a Content-signal line and a policy comment block.
What are the three signals?search covers indexing and returning links and excerpts. ai-input covers feeding content into AI models, including retrieval augmented generation and live generative answers. ai-train covers training or fine-tuning models. Each is set to yes or no.
What is the default Cloudflare applied?search=yes and ai-train=no. Cloudflare deliberately left ai-input neutral, because it is the one signal where publishers genuinely disagree and a vendor default would be guessing at strategy.
Does Google honor content signals?No. Google has stated the directive has no effect on Google whatsoever. Googlebot reads standard robots.txt rules and Google's own tokens, and ignores third-party extensions. Content signals reach other retrieval operators, not Google.
Is it legally binding?Not on its own. Cloudflare's documentation states robots.txt compliance is voluntary. The policy is framed as an express reservation of rights under Article 4 of EU Directive 2019/790, which is a documented position rather than an enforcement mechanism.
Will setting ai-input=yes get me into AI Overviews?No. Google does not read the signal, so it changes nothing on Google surfaces. It states your position to operators that do read robots.txt usage signals, which is a smaller but growing share of where buyers get answers.
Should a B2B SaaS company set ai-input to yes?In almost every case, yes. If being named when an engine shortlists vendors in your category is an acquisition event, declaring that use costs nothing and removes an ambiguity that currently resolves against you.
How do I change it if Cloudflare manages the file?Change it in the Cloudflare dashboard, not by editing robots.txt directly. Managed files regenerate, and a hand-edit is reverted at the next regeneration without notifying anyone.
Does blocking a crawler and setting a signal do the same thing?No. Blocking stops the fetch and is enforced by your infrastructure. A signal permits the fetch and states what the operator may do afterward, enforced only by the operator's willingness to comply.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

TT
Tyler TruffiMANAGING PARTNER, SOMETHING INC.

Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.