Something Inc.Schedule a free consultation
TECHNICAL SEO

We told clients to ship this file. New data says it doesn't help.

Two independent studies just scanned a combined 337,894 domains for llms.txt. Neither found a citation bump. Here's what we're telling clients now instead.

TTTyler TruffiManaging Partner · JUL 24, 2026 · 8 MIN READ
KEY TAKEAWAYTwo independent 2026 studies, covering a combined 337,894 domains, found no measurable citation lift from having an llms.txt file. The file isn't the thing that was working. Machine access, the thing the file was supposed to signal, still is.

If you've been reading our content this year, you've seen the format recommended more than once. That's not a footnote I want to bury. It's the actual reason this piece exists: a correction only counts if it's as visible as the original claim, and this one is going at the top, not somewhere it's easy to skip past.

I've spent a chunk of this year telling clients to ship an llms.txt file. So has half the GEO industry. We wrote the guide. We gave you the exact format. New data says it didn't do what we said it would.

That's not a fun sentence to write. It's also the whole job. You don't get to publish the guide and then go quiet when someone runs the numbers and the numbers don't agree with you. So let's go through what actually landed this month, and what we're telling clients to do differently starting now.

The confession

A prototype is not a product. And it turns out a text file is not a strategy, either. Here's the thing about llms.txt: it always sounded almost too tidy. A plain-text map that tells AI crawlers who you are and where your best content lives, sitting quietly at the root of your domain, doing the work robots.txt does for search bots. It was a clean idea. Clean ideas get adopted fast, and get tested even faster.

That's exactly what happened. And the test results are in.

I want to be specific about who "we" is here, because it's not some faceless corner of the industry. It's us. It's most of the vendor blogs you follow. It's the newsletter writers who covered the format as a quick, high-leverage win the moment it started circulating. Nobody ran a controlled test before recommending it, because the theory was clean enough that it felt like it didn't need one. That's exactly the kind of advice that deserves the most scrutiny, and exactly the kind that usually gets the least, because it feels too obviously correct to bother checking.

What the 300,000-domain study found

SE Ranking ran the biggest version of this test: 300,000 domains, checked for the presence of an llms.txt file and then cross-referenced against actual AI citation frequency. Search Engine Journal covered the results in July. The file showed up on 10.13% of the domains scanned. Among the sites that had it, there was no clear, consistent citation-frequency lift compared to sites that didn't. Glenn Gabe flagged the study to his own audience within hours of it landing, which is usually a sign an SEO finding is about to become a talking point, not a footnote.

300,000
domains scanned in the SE Ranking study
10.13%
of domains had an llms.txt file at all
0
measurable citation-frequency advantage for having one

A second study says the same thing

One study is a data point. Two independent studies landing on the same conclusion is a pattern. Trakkr Research ran its own version, 37,894 domains, titled bluntly: "The llms.txt Effect: Zero Citation Advantage." Same design, same negative result, different team, different sample. That's the kind of replication you don't get to wave away.

STUDYDOMAINS SCANNEDCITATION LIFT FOUND
SE Ranking (Jul 2026)300,000None measurable
Trakkr Research (2026)37,894None measurable
Rankability adoption data (Jun 2026)Top 1,000 sites8.7% adoption; not a citation study

That last row matters because it shows adoption climbing while effectiveness stays flat. Rankability separately measured 8.7% adoption among the top 1,000 sites as of June. People are shipping the file. It's just not doing the thing they think it's doing.

Notice, too, what these two studies didn't find. Neither one found that llms.txt actively hurts a site's citation rate. That distinction matters, because the honest reading of "zero measurable lift" is closer to "inert" than "harmful." A file that does nothing is a much smaller problem than a file that actively misdirects a crawler, and it means nobody needs to panic-delete anything this afternoon. The panic-worthy part isn't the file sitting on your server. It's the hours of strategy time that got spent treating it as the centerpiece of a GEO rollout instead of a footnote.

"llms.txt is complete BS. Google's crawler can browse websites, run JS, and read everything fine. Why do you think AI crawlers won't?" — Pieter Levels, in a widely-shared post that aged better than most of us wanted it to.

What llms.txt was actually supposed to fix

Here's where I have to be honest about what we got wrong, and what we didn't. The theory behind llms.txt was never crazy: AI crawlers have limited time and budget per site, so give them a map. Point straight at your best pages. Skip the maze. That's a real problem. It's just not the problem llms.txt turned out to solve.

A Hacker News thread put it more bluntly than any vendor blog would: "LLMs are not reading llms.txt nor AGENTS.md files." The modern crawler behind ChatGPT, Perplexity, and Claude doesn't need a hand-written index. It renders the page, follows the links, and builds its own map, the same way a determined human would. A text file sitting off to the side asking politely to be noticed just isn't part of that loop in the way early advocates assumed.

There's a version of this mistake that goes back further than llms.txt. Every time a new machine-facing standard shows up, a chunk of the industry treats publishing the file as equivalent to doing the underlying work the file was supposed to represent. It happened with meta keyword tags. It happened, to a lesser degree, with some of the more exotic schema types that never got broad parser support. llms.txt is the 2026 version of the same instinct: a static declaration standing in for the dynamic, harder-to-fake reality of whether a crawler can actually reach, parse, and trust your content.

What still works

This is not an argument that machine access doesn't matter. It's the opposite. The signals that actually decide citation, extractable structure, demonstrated authority, and machine access, are all still real. llms.txt was supposed to be a shortcut to the third one. It wasn't a bad instinct. It was a shortcut that didn't survive contact with how these crawlers actually behave.

1Crawlability, the boring kind, still mattersConfirm GPTBot, PerplexityBot, and ClaudeBot aren't blocked in robots.txt, aren't rate-limited into timeouts, and aren't stuck behind a JavaScript wall. This is the access layer that was never optional.
2Structure still gets you extractedClean headings, direct answers near the top, and tables the model can lift without paraphrasing. This is what the file was trying to shortcut, and there's no shortcut for it.
3Corroboration still gets you citedBeing named on the third-party sources engines already trust does more for citation rate than any file at the root of your domain.

What to do instead

I'll also say the quiet part about why this took until July to surface clearly: nobody wanted to be the one running the negative study. Vendors selling GEO tooling had every incentive to keep recommending a quick, sellable checklist item. Agencies, including us, had a guide already published and a client deck already built around it. The correction only shows up once someone runs the numbers with no horse in the race, which is exactly what SE Ranking and Trakkr Research both did, independently, within weeks of each other.

So here's what changes on our end, starting with this piece. We're not telling clients to rip out an existing llms.txt file; it's not actively harmful, and it costs nothing to leave in place. We are telling clients to stop treating it as a checkbox that earns you anything, and to stop leading a GEO audit with it. The afternoon it used to take gets redirected: confirm crawler access is clean, run a real technical audit on how the pages actually render for a crawler that can't execute your JavaScript, and fix the structure of your three highest-value pages so an engine can lift a fact from them without guessing.

One more thing worth saying plainly: this is exactly the kind of correction this column exists to make. A guide we published in good faith, based on a reasonable theory and the industry consensus at the time, met new data that didn't support it, and the right move is saying so in public rather than quietly stopping the recommendation and hoping nobody notices the gap between what we said in June and what we're doing in August.

You're not behind for having shipped an llms.txt file this year. Most of the industry did the same thing on the same theory, ourselves included, and it was a reasonable bet before the data came in. What would be behind is still leading a client GEO plan with it in August, after two independent studies covering 337,894 domains came back with the same answer. Cross it off the top of the list. Put crawler access and page structure back at the top, because that engagement work is where the actual lift keeps coming from, and it's the same lesson our headless-site clients already learned the hard way with plain search crawlers. We ran the exact version of this recalibration on a cybersecurity platform this year: dropped the llms.txt line item, doubled down on machine access and structure, and the citation rate moved anyway. The file was never the mechanism. It just felt like one.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

TT
Tyler TruffiMANAGING PARTNER, SOMETHING INC.

Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.