Something Inc.LoginSchedule a free consultation
TECHNICAL SEO

PDF SEO in 2026: Google is quietly dropping your PDFs

Top-performing PDFs went to zero impressions in mid-August, including IRS and state tax forms. Google has not explained it. If your best assets are PDFs, treat this as the warning shot.

TECHNICAL SEOREPORTED AUG 12 TO AUG 26, 2026
AUG 18
date the broader PDF impression dip began, per practitioner reports
AUG 12
date the earliest reported PDF drop to zero impressions started
10K
impressions the two PDFs one marketer lost had produced over three months
THE SHORT VERSIONMultiple SEOs reported PDFs falling out of Google results starting in mid-August 2026, including federal and state tax forms. Google's John Mueller acknowledged the reports without confirming a cause. Nobody outside Google knows whether this is a bug, a deliberate quality change, or fallout from the August 2026 spam update. All three lead to the same decision for enterprise teams.
TL;DR · 60 SECONDSPDF SEO has been a tolerated exception for two decades: Google indexed PDFs, ranked them, and most teams never questioned parking a whitepaper or a spec sheet in one. Starting around August 12 and widening around August 18, practitioners began reporting PDFs going to zero impressions in Search Console, with IRS and New York State tax forms among the visible casualties. There is no confirmed cause. What there is, is a clear reason to stop treating a PDF as a permanent, indexable, citable asset. The fix is not a PDF optimization checklist. It is publishing the content as HTML and keeping the PDF as the download, not the destination.

Every enterprise site we audit has a PDF problem it does not know about. Spec sheets, compliance documents, gated research, product one-pagers, board-approved regulatory filings. They sit at /assets/whatever-final-v3.pdf, they pick up links and impressions, and nobody has looked at them in two years because they have always just worked. Mid-August was the first serious sign that always is doing a lot of work in that sentence.

What actually happened to PDFs in Google

The thread started with Savanna Gray flagging on LinkedIn that PDFs she tracked had stopped appearing in Google results. Lily Ray picked it up and pushed it wider on LinkedIn and X, and more practitioners piled in with matching observations. Tamara Helgren shared Search Console data. Search Engine Roundtable collected the reports on August 26. Daniel Deceuster, VP of Marketing at Zion HealthShare, gave the most concrete account of the damage: his top PDF went to zero impressions starting August 18, his second best performer went to zero on August 12, and between them those two files had produced roughly 10,000 impressions and hundreds of clicks over the previous three months.

The examples people surfaced are the tell. These were not thin, spammy files. They were IRS forms including the W-9, W-4 and 940, New York State tax forms, Connecticut employment forms, and, at the other end of the spectrum, printable coloring books. Government tax forms are about as close to canonical, authoritative, high-intent documents as the open web produces. If those are the files losing visibility, a quality-based explanation gets harder to defend.

WHAT WAS OBSERVEDDETAILWHERE IT CAME FROM
Earliest dropsIndividual PDFs to zero impressions starting August 12Search Console data shared publicly by Daniel Deceuster, Zion HealthShare
Broader dipWider reports of PDF impressions falling from August 18Practitioner reports collected by Search Engine Roundtable, August 26
Overlapping eventGoogle's August 2026 spam update ran August 18 to August 21Google Search Status Dashboard
Named examplesIRS W-9, W-4 and 940, New York State tax forms, Connecticut employment formsReports from Savanna Gray, Lily Ray and others on LinkedIn and X
Google's responseJohn Mueller said he would look into it. No cause confirmed.Mueller, on Bluesky

Read the timing column carefully, because it is the part most coverage glosses. The earliest confirmed drop predates the spam update by six days. That matters. If every reported drop clustered on August 18, the obvious read would be spam update collateral and the obvious advice would be to wait for the dust to settle. A drop that starts on August 12 and widens on August 18 looks more like two things happening near each other, or one slow change that the update happened to accelerate. Either way, waiting is a worse plan than it looks.

Google acknowledging a report is not Google confirming a bug. Plan for the version where this is intentional.

Mueller's reply, offered on Bluesky, was that he would take a look and thanks for passing it on. That is a genuinely useful signal about whether Google is aware. It is not a signal about whether the behavior is going to be reversed. We have watched enough of these cycles to know the honest planning assumption: treat an acknowledged report as roughly a coin flip between fix and feature, and make the decision that survives both outcomes.

Why PDF SEO was already on borrowed time

Strip out the August drama and PDFs were already the weakest surface on most enterprise sites. Google has indexed PDFs for years, which built a false sense of parity with HTML, but the format loses on nearly every dimension that decides modern visibility.

A PDF has no meaningful internal linking, so it absorbs authority and passes almost none of it back into your site architecture. It carries no schema, which means no structured data for products, FAQs, authors or organizations. It has no title tag or meta description in the sense a search engine or a summarizer wants, only document metadata that half the design tools populate incorrectly. It cannot be updated without changing the file, which means either a stale URL or a broken one. It renders badly on phones. It reports poorly in analytics. It is invisible to most on-site search. And critically, it is a single monolith: a 40-page PDF is one URL competing for every topic inside it, where the same content as HTML would be a dozen pages each competing for one.

CAPABILITYHTML PAGEPDF
Structured data and schemaFull supportNone
Internal links out to your siteNativeRare and usually broken
Updatable without a new URLYesOnly by replacing the file
One URL per topicYesOne URL for the whole document
Mobile renderingResponsivePinch and zoom
Machine-readable page structureHeadings, lists, tablesDepends entirely on export settings

None of that is new. What is new is that the tolerated exception stopped being tolerated for a couple of weeks, and everyone got to see what their traffic looks like without it. If a format only works while a search engine is being generous about it, that format is a dependency, not a strategy.

AI engines have their own PDF problem

There is a second reason to care, and for most of our clients it now outweighs the first. Generative engines have to extract a self-contained, quotable claim from a page in order to cite it. That is the whole mechanic, and it is why content structure changes citation rate so sharply. PDFs are hostile to extraction in a way HTML is not.

A PDF is a description of where ink goes on a page, not a description of what the content means. Headings are frequently just larger text with no semantic marker. Multi-column layouts interleave when parsed linearly. Tables become runs of loose numbers. Charts are images with no alternative text. A retrieval system can often get something out of a PDF, but it gets less, less reliably, and with a worse sense of what the document is actually claiming. When the same content exists as HTML somewhere else on the web, the HTML version wins the citation on structure alone.

The crawler side compounds it. Fetch behavior differs meaningfully across AI user agents, and our work on what AI crawlers actually fetch found that plenty of them are far more selective than Googlebot about non-HTML resources. Bandwidth is not free for them either. A large binary that requires a parsing step is exactly the kind of resource a retrieval pipeline deprioritizes when it has an HTML alternative. So a company that ships its best research as a gated PDF is often absent from AI answers on its own topic while a competitor's thinner blog post gets named.

Which of your PDFs are actually at risk

Not every PDF is a problem, and a blanket migration is a waste of a quarter. Sort your library into four buckets and only two of them need work.

1Content that should never have been a PDFWhitepapers, research reports, buyer guides, blog-style thought leadership. These earn impressions, links and citations, and every one of those benefits is capped by the format. Migrate to HTML first, keep the PDF as an optional download at the bottom.
2Reference documents people search for by nameSpec sheets, compatibility matrices, pricing tables, integration lists, compliance summaries. These get searched with real intent and answered by AI engines constantly. They belong on an HTML page with a table an engine can parse, not inside a file it has to guess at.
3Documents that must stay documentsSigned contracts, filed regulatory submissions, forms designed to be printed and mailed, anything where the fixed layout is the point. Leave these as PDFs and give each one an HTML landing page that describes it, links to it, and carries the schema and the copy that earns the ranking.
4Dead weightEvent decks from 2023, superseded datasheets, duplicate versions with -final-v2 in the filename. Redirect them to the current HTML equivalent and stop maintaining them. Most enterprise libraries are a third dead weight and nobody has ever counted.

The triage question that sorts most files in about five seconds: if a person read only this document and never visited your site, did you win or lose? Whitepapers and buyer guides lose, because the whole point is to pull the reader into your funnel and a PDF is a dead end with no links. A signed compliance attestation wins, because being a faithful, unchangeable record is exactly its job.

The PDF SEO migration that takes a week

This is a smaller project than it sounds, and it is one of the highest-return weeks a technical team can spend. Pull every PDF URL out of Search Console with impressions or clicks in the last twelve months, and every PDF with referring domains out of your backlink tool. That intersection is your real library. For most mid-market sites it is between 20 and 60 files, not the 400 sitting in the CMS.

For each file in buckets one and two, publish the content as a real HTML page with proper headings, a table of contents, schema, an author byline, and internal links out to the relevant service and product pages. Keep the PDF alive at its original URL so existing links do not break, and add a canonical or a redirect depending on whether the file still needs to be downloadable. Then link the HTML page from the PDF's own first page, so the assets already circulating in inboxes and Slack channels route people back to something indexable.

The PDF triage query set● LIVE
Search Console filter: Page contains .pdf, last 12 months, sort by impressions
Backlink tool filter: target URL contains .pdf, sort by referring domains
Site crawl filter: content-type application/pdf, export all
Cross-reference the three, keep anything appearing in two of them
Everything else is dead weight, redirect it and move on

Two details teams get wrong. First, do not delete the PDFs. Deleting a file with 40 referring domains to fix an indexing concern trades a small problem for a large one. Second, do not make the HTML page a summary that teases the PDF. If the HTML version withholds the substance, engines will not cite it and readers will not link it, and you have rebuilt the same dead end with extra steps. Publish the whole thing.

Do this before your next quarter closes

Open Search Console, filter Page for .pdf, and set the date range to the last 28 days compared against the previous 28. If your PDF impressions fell off a cliff around August 12 or August 18, you are in the affected group and you now have a documented, dated business case for the migration budget you have been quietly asking for. If they did not fall, you got a free warning and a window to act in without an emergency attached.

Then pick the three PDFs with the most referring domains and publish HTML versions of them this month. Not all 40. Three. That gives you a before-and-after on the same domain with the same authority, and it is a far more persuasive internal argument than any external case study, including ours. The teams that ran this migration in 2024 for accessibility reasons are the ones sitting comfortably this week, which is usually how these things go.

If the library is large enough that triage alone is a project, that is what a technical SEO program or a full content marketing engagement is for, and it is a recurring finding for fintech and other regulated categories where half the useful content on the site is legally required to exist as a document. The point is not that PDFs are bad. The point is that a PDF should be a copy of your content, never the only version of it. Google spent two weeks in August demonstrating why, and it did not owe anyone an explanation first.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

JB
Josh BernsteinMANAGING PARTNER, SOMETHING INC.

Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.