Something Inc.Schedule a free consultation
TECHNICAL SEO

97% of llms.txt Files Are Never Read, Ahrefs Finds

Ahrefs analyzed 137,210 domains and found that 97% of published llms.txt files got zero requests in a full month. Here's what the traffic that did land tells you about who's actually reading the file.

TTTyler TruffiManaging Partner · AUG 3, 2026 · 9 MIN READ

Ahrefs just ran the largest llms.txt readership study to date, and the number that matters isn't adoption. It's this: across 137,210 domains, 97% of published llms.txt files received zero requests in May 2026. Not low traffic. Zero. If you spent an afternoon this year drafting an llms.txt file as a quick generative engine optimization win, the data says almost nobody, human or bot, ever opened it.

TL;DR · 60 SECONDSAhrefs analyzed 137,210 domains using its own Web Analytics and Bot Analytics products, examining server logs and live traffic during May 2026 (published June 15, 2026 by Louise Linehan and Xibeijia Guan). 28% of domains — 38,360 sites — publish a valid llms.txt file. 97% of those files never got a single request. Of the requests that did land, 96% came from bots, but named AI tools accounted for only 19.5% of that traffic and industry auditing tools for 12%. Zero bots even searched for llms.txt files on domains where none existed. Ahrefs' own read: the file's actual readership is agentic coding tools, not AI search or answer engines.

The 137,210-Domain Study Behind llms.txt's 97% Problem

Ahrefs pulled this from real infrastructure, not a survey. The team used its own Web Analytics and Bot Analytics products to examine server logs and live traffic across 137,210 domains during May 2026, then published the results June 15, 2026 in Ahrefs' study, written by Louise Linehan and Xibeijia Guan. That's a meaningfully different method than the adoption counts that have circulated all year — this one watched what actually requested the file, not just whether the file existed.

Adoption itself isn't rare anymore. 28% of the domains scanned — 38,360 sites — publish a valid llms.txt file, with the file itself having become one of the most recommended quick wins on generative engine optimization checklists this year. But adoption and readership turned out to be two entirely different questions. Of those 38,360 published files, 97% received zero requests during the entire month Ahrefs measured. Not "low traffic." Zero traffic. The file sat at the root of the domain, doing nothing, for the overwhelming majority of the sites that bothered to publish one.

137,210
domains Ahrefs analyzed in May 2026
28%
of domains publish a valid llms.txt file
97%
of published llms.txt files got zero requests

Compare that to how teams have been pricing the work. A well-run llms.txt rollout gets treated as a checklist line item, an hour or two to draft the file, point it at your best pages, and ship it. That math only holds if the file does something on the other end. Ahrefs' numbers say it usually doesn't. Across 38,360 published files, 97% sat unread for the entire month measured, which means the actual return on that hour, for most sites, rounds to zero.

Who Actually Requests an llms.txt File

The 3% of files that did get a request tell you more than the 97% that didn't. Ahrefs broke down that traffic, and it undercuts the premise most teams are working from. 96% of requests came from bots — expected, since a plain-text file at the root of a domain isn't a page a human stumbles onto while browsing. But split that 96% further and the picture changes. Named AI tools, the ChatGPTs, Perplexitys, and Claudes most teams assume they're building the file for, accounted for only 19.5% of requests. Industry auditing tools, the GEO and AEO platforms and researchers checking whether the format is catching on, accounted for 12%.

All bot traffic96%
Named AI tools specifically19.5%
Industry/GEO auditing tools12%
Chrome Lighthouse audits0.1%

Share of the fetches that actually happened, by source (Ahrefs, May 2026)

Chrome Lighthouse audits, which ping llms.txt as part of a routine accessibility and performance check, generated roughly 1 in 1,000 of the fetches that happened at all — a rounding error, but a telling one. A chunk of the small amount of traffic you do see isn't intentional AI retrieval. It's tooling side effects.

"If you publish an llms.txt file today, the most likely outcome by far is that nothing ever fetches it." — Ahrefs, June 2026

Ahrefs went further and checked whether AI bots were even speculatively probing for the file on domains that didn't have one, the kind of behavior you'd expect if crawlers routinely checked llms.txt the way they check robots.txt. Zero bots did. Not rarely. Zero. That's a stronger result than "AI engines don't prioritize it" — it says AI retrieval bots aren't looking for the file at all, on sites with one or without one. Ahrefs' own conclusion: actual AI retrieval bots show minimal interest in llms.txt, and the file's primary real readership is agentic coding tools, not AI search or answer engines. That reframes what the format is actually for. It was pitched as a GEO tactic, a way to get cited more in AI answers. The traffic says its real audience, thin as it is, is developer tooling.

That lines up with what we found when we checked citation lift directly. Two prior studies covering 337,894 domains found no measurable citation lift from having an llms.txt file at all. Ahrefs' server-log data explains why. You can't get a citation lift from a file that's rarely fetched by the systems that would need to fetch it in order to cite you.

Even Without llms.txt Blocking, Most Sites Stay Invisible

llms.txt's readership problem isn't the only invisibility problem sitting in this space. Website Auditor's AI Visibility Index, from Armstrong HoldCo LLC, published a second edition on August 2, 2026 (the first edition ran July 31, 2026), covering 531 audits across 458 domains in 38 sectors. The topline finding: only 8.9% of sites block at least one AI crawler via robots.txt. Most sites aren't the access problem GEO vendors love to warn you about. And yet 94.8% of the audited sites never appear in an AI assistant answer at all, based on 193 sites and 5,978 answers checked across ChatGPT, Claude, Gemini, and Perplexity between July 23 and August 2. Gemini had the highest citation rate of the four, and it topped out at 2.9%.

Read that sample size honestly. The study's own authors caution it's self-selected, not a random sample, skewing toward small and mid-size businesses already working on AI visibility. The exact percentages are directional, not universal, and a fully random sample of the web could land differently. The pattern underneath them is still worth sitting with: 8.9% blocked, 94.8% invisible anyway.

METRICWEBSITE AUDITOR AI VISIBILITY INDEX (2ND ED., AUG 2, 2026)
Sites blocking at least one AI crawler via robots.txt8.9%
Sites never appearing in an AI assistant answer94.8% (of 193 sites / 5,978 answers)
Highest single-engine citation rateGemini, 2.9%
Sites publishing LocalBusiness schema19.3%
Sites publishing any schema.org markup54.6%

Put the two studies side by side and the throughline is obvious. Publishing llms.txt is being fetched by almost nobody. Not blocking crawlers doesn't get you cited either — roughly 91% of the Website Auditor sample already leaves AI crawlers unblocked, and 94.8% is still invisible in AI answers anyway. Both "publish the file" and "don't block the bots" turn out to be necessary conditions and nowhere close to sufficient ones. Neither is the lever. They're the floor.

What Actually Moves AI Citations Instead of llms.txt

If access and file publication aren't the lever, what is? The Website Auditor data hints at the answer from a different angle: only 54.6% of audited sites publish any schema.org markup at all, and only 19.3% publish LocalBusiness schema specifically. That's a structure gap, not an access gap, and structure is one of the few variables in this entire picture that a site controls end to end.

01Extractable structureClean headings, direct answers placed near the top, and tables an engine can lift without paraphrasing. No file at the root of your domain substitutes for this.
02Demonstrated authorityBeing named on the third-party sources an engine already trusts does more for citation odds than anything self-published, llms.txt included.
03Verified machine accessAn unblocked robots.txt line isn't proof a crawler can actually render your content. JavaScript-heavy and headless sites routinely fail this even when nothing is technically blocked.

None of these three require an llms.txt file. All three require real engineering and content work, which is exactly why they're less fun to write about than a text file you can generate in an afternoon. There's also a cost side to crawler access that most llms.txt advocates skip entirely: AI crawlers hit your infrastructure whether or not they ever cite you, and the economics of that traffic are a separate decision from whether you're optimizing for citation in the first place.

There's a broader lesson here about how the GEO industry keeps treating standards. A new format shows up, it's easy to explain in a client deck, and it gets adopted faster than anyone tests whether it's actually load-bearing. Ahrefs' server logs are the closest thing this space has had to a ground-truth check on that instinct, and the check came back negative. That doesn't make the underlying intuition behind llms.txt wrong, giving crawlers a map is still a reasonable idea, it makes the current implementation of that idea a low-priority one, not a foundational one.

What to Do With Your llms.txt File Now

Leave the file up if you already have one. The data doesn't say it hurts anything, only that it isn't doing the job it was sold as. What changes is where you spend the next audit hour.

KEY TAKEAWAYllms.txt costs you an afternoon and nothing else. The next audit hour is worth more spent on schema markup, extractable structure, and verified crawler access than on tuning a file that 97% of the time nobody requests.

This matters most for B2B SaaS sites, where buyer research increasingly runs through an AI assistant before it ever reaches a search results page, and where the schema and structure gap Website Auditor found, only 54.6% publishing any markup at all, is the more common failure mode than crawler blocking.

Run the audit that actually matters: confirm GPTBot, ClaudeBot, and PerplexityBot can render your key pages without hitting a JavaScript wall, add or fix schema on the pages driving pipeline, and rewrite your three highest-value pages so an engine can lift a direct answer without guessing at structure. That's the actual mechanics behind generative engine optimization work, and it's where both studies above point once you stop treating llms.txt as the finish line. If you already have the file, don't touch it. If you don't, stop putting it at the top of the list. Put the audit above it, this week.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

TT
Tyler TruffiMANAGING PARTNER, SOMETHING INC.

Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.