Here's a thing that happened today, Aug 4, 2026. Mark Williams-Cook, who runs Candour and built AlsoAsked, published a post about a file called cats.txt. It's not a real standard. He made it up. It describes, in a straight face, some office cats. And then he sat back and watched the exact same chain of "proof" that gets cited for llms.txt happen to a document about cats.
The experiment: cats.txt
You know the pitch by now. Someone tells you llms.txt is worth doing because AI crawlers fetch it, because Google indexes it, because you can ask ChatGPT about your site and it name-drops the file, because "a major AI company recommended it." Four proof points, repeated across a hundred agency decks and vendor blog posts this year.
Williams-Cook's move was almost embarrassingly simple, which is exactly why it lands. He published cats.txt: a plain-text file, in the same format as llms.txt, describing a handful of fictional office cats. Then he waited. PerplexityBot showed up. GPTBot showed up. ClaudeBot showed up. Google indexed the page. He asked ChatGPT about cats.txt directly, and it answered helpfully, the same tone it uses when someone asks about robots.txt or llms.txt. Every single proof point people cite for llms.txt working, cats.txt hit too. And cats.txt is about cats. It does nothing. It was never supposed to do anything.
Why crawling, indexing, and citation don't prove llms.txt works
Take the four proof points one at a time, because each one collapses the same way once you ask what it would take for the proof point to fail.
The convergence problem
Williams-Cook's name for this is the convergence problem, and it's worth sitting with because it's a sharper diagnosis than "AI hallucinates." A model doesn't independently evaluate whether llms.txt improves citation odds. It has absorbed thousands of confident blog posts, LinkedIn takes, and vendor pitches that already assume the answer, and it reflects that consensus back at you when you ask. If the consensus is wrong, unverified, or built entirely on the four hollow proof points above, the model repeats the wrong answer with total confidence, because confidence was never the thing it was measuring. It's an echo, dressed up as a verdict, and it is very good at sounding like a verdict.
“"Being in the index is a statement that a URL exists and contains words. It is not a verdict on truth, usefulness or sanity." — Mark Williams-Cook, Aug 4, 2026.”
It's worth being precise about what cats.txt does and doesn't demonstrate, because the temptation is to overreach in the other direction now and declare the whole exercise of machine-readability files worthless. That's not the finding. The finding is narrower and more useful than that: the specific four proof points people cite for llms.txt's value don't discriminate between something real and something fabricated, so they can't be used as evidence either way. That leaves the actual question wide open, which is uncomfortable, but it's a more honest place to stand than a comforting stack of hollow proof points.
This is where the GEO industry, ours included, has to be honest about its own supply chain. A claim gets published somewhere with a number attached. It gets summarized in a LinkedIn post. The LinkedIn post gets summarized in a listicle. The listicle gets crawled. Six months later, an AI engine repeating that claim back to a user, confidently, with no hedge, feels like independent confirmation. It isn't. It's the same claim, laundered through enough intermediate steps that its original shakiness is no longer visible. cats.txt compresses that whole laundering process into one afternoon, which is what makes it such a clean demonstration.
This isn't really about llms.txt
We've written before about the actual measured data on llms.txt: SE Ranking's 300,000-domain study and Trakkr Research's 37,894-domain study both found zero measurable citation lift from publishing the file. Ahrefs went further, tracking 137,210 domains and finding 97% of published llms.txt files got zero requests in a full month. Those are real, sourced, falsifiable numbers, gathered by actually measuring outcomes rather than checking whether a crawler showed up once. That's the difference cats.txt is pointing at: real evidence has a mechanism by which it could have come back negative. Crawl visits and index inclusion can't come back negative for anything, including cats.txt, which is why they were never evidence in the first place.
The reason this deserves a piece of its own, separate from the llms.txt data pieces, is that the pattern generalizes far past one text file. Duane Forrester ran an open, deliberately skeptic-solicited survey this summer asking whether GEO and AI-visibility platforms deliver real value at all, precisely because he suspected the category was running on the same kind of soft, self-confirming proof. Whether a GEO tool is actually proven is the same question as whether llms.txt is actually proven, asked about a different product category. Vendor case studies that show a mention-rate number going up, with no controlled comparison and no counterfactual, are cats.txt with a dashboard attached. A screenshot of ChatGPT praising a client's visibility strategy is cats.txt with a client logo attached. The tell is always the same: does the proof point have a way to fail? If a fake file passes every test a real one does, the tests weren't testing anything.
| CLAIM TYPE | COULD THIS PROOF POINT FAIL FOR SOMETHING FAKE? | VERDICT |
|---|---|---|
| "AI crawlers fetch it" | No — cats.txt got fetched too | Not evidence |
| "Google indexed it" | No — cats.txt is indexed | Not evidence |
| "ChatGPT mentions it" | No — reflects current discourse, not ground truth | Not evidence |
| "Measured citation lift across N domains, pre/post" | Yes — a real study can return a null result | Actual evidence |
| "Independently reproduced by a second dataset" | Yes — a second team could fail to replicate it | Actual evidence |
How to actually read a GEO claim from now on
None of this is an argument for cynicism about GEO as a discipline, or an excuse to stop measuring anything because measurement is hard. It's the opposite. It's an argument for spending your skepticism on the right target. Don't doubt that AI search visibility matters; the citation and mention-rate data we've published all year is real, sourced, and reproducible. Doubt any specific claim that can't tell you what a negative result would have looked like.
Something Inc. runs generative engine optimization work on the premise that most of it is measurable, falsifiable, and worth doing properly, which is exactly why the sloppy version of the discipline is worth calling out in public. The technical-access side of this matters too, separate from the evidence-quality problem: a page that's genuinely invisible to AI crawlers because it never renders past a JavaScript shell has a real, fixable, falsifiable problem. A page that publishes llms.txt and expects that alone to move a citation number does not. Keep those two categories of problem straight and you'll spend your GEO budget on the one that actually moves.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.