Every enterprise site we audit has a PDF problem it does not know about. Spec sheets, compliance documents, gated research, product one-pagers, board-approved regulatory filings. They sit at /assets/whatever-final-v3.pdf, they pick up links and impressions, and nobody has looked at them in two years because they have always just worked. Mid-August was the first serious sign that always is doing a lot of work in that sentence.
What actually happened to PDFs in Google
The thread started with Savanna Gray flagging on LinkedIn that PDFs she tracked had stopped appearing in Google results. Lily Ray picked it up and pushed it wider on LinkedIn and X, and more practitioners piled in with matching observations. Tamara Helgren shared Search Console data. Search Engine Roundtable collected the reports on August 26. Daniel Deceuster, VP of Marketing at Zion HealthShare, gave the most concrete account of the damage: his top PDF went to zero impressions starting August 18, his second best performer went to zero on August 12, and between them those two files had produced roughly 10,000 impressions and hundreds of clicks over the previous three months.
The examples people surfaced are the tell. These were not thin, spammy files. They were IRS forms including the W-9, W-4 and 940, New York State tax forms, Connecticut employment forms, and, at the other end of the spectrum, printable coloring books. Government tax forms are about as close to canonical, authoritative, high-intent documents as the open web produces. If those are the files losing visibility, a quality-based explanation gets harder to defend.
| WHAT WAS OBSERVED | DETAIL | WHERE IT CAME FROM |
|---|---|---|
| Earliest drops | Individual PDFs to zero impressions starting August 12 | Search Console data shared publicly by Daniel Deceuster, Zion HealthShare |
| Broader dip | Wider reports of PDF impressions falling from August 18 | Practitioner reports collected by Search Engine Roundtable, August 26 |
| Overlapping event | Google's August 2026 spam update ran August 18 to August 21 | Google Search Status Dashboard |
| Named examples | IRS W-9, W-4 and 940, New York State tax forms, Connecticut employment forms | Reports from Savanna Gray, Lily Ray and others on LinkedIn and X |
| Google's response | John Mueller said he would look into it. No cause confirmed. | Mueller, on Bluesky |
Read the timing column carefully, because it is the part most coverage glosses. The earliest confirmed drop predates the spam update by six days. That matters. If every reported drop clustered on August 18, the obvious read would be spam update collateral and the obvious advice would be to wait for the dust to settle. A drop that starts on August 12 and widens on August 18 looks more like two things happening near each other, or one slow change that the update happened to accelerate. Either way, waiting is a worse plan than it looks.
“Google acknowledging a report is not Google confirming a bug. Plan for the version where this is intentional.”
Mueller's reply, offered on Bluesky, was that he would take a look and thanks for passing it on. That is a genuinely useful signal about whether Google is aware. It is not a signal about whether the behavior is going to be reversed. We have watched enough of these cycles to know the honest planning assumption: treat an acknowledged report as roughly a coin flip between fix and feature, and make the decision that survives both outcomes.
Why PDF SEO was already on borrowed time
Strip out the August drama and PDFs were already the weakest surface on most enterprise sites. Google has indexed PDFs for years, which built a false sense of parity with HTML, but the format loses on nearly every dimension that decides modern visibility.
A PDF has no meaningful internal linking, so it absorbs authority and passes almost none of it back into your site architecture. It carries no schema, which means no structured data for products, FAQs, authors or organizations. It has no title tag or meta description in the sense a search engine or a summarizer wants, only document metadata that half the design tools populate incorrectly. It cannot be updated without changing the file, which means either a stale URL or a broken one. It renders badly on phones. It reports poorly in analytics. It is invisible to most on-site search. And critically, it is a single monolith: a 40-page PDF is one URL competing for every topic inside it, where the same content as HTML would be a dozen pages each competing for one.
| CAPABILITY | HTML PAGE | |
|---|---|---|
| Structured data and schema | Full support | None |
| Internal links out to your site | Native | Rare and usually broken |
| Updatable without a new URL | Yes | Only by replacing the file |
| One URL per topic | Yes | One URL for the whole document |
| Mobile rendering | Responsive | Pinch and zoom |
| Machine-readable page structure | Headings, lists, tables | Depends entirely on export settings |
None of that is new. What is new is that the tolerated exception stopped being tolerated for a couple of weeks, and everyone got to see what their traffic looks like without it. If a format only works while a search engine is being generous about it, that format is a dependency, not a strategy.
AI engines have their own PDF problem
There is a second reason to care, and for most of our clients it now outweighs the first. Generative engines have to extract a self-contained, quotable claim from a page in order to cite it. That is the whole mechanic, and it is why content structure changes citation rate so sharply. PDFs are hostile to extraction in a way HTML is not.
A PDF is a description of where ink goes on a page, not a description of what the content means. Headings are frequently just larger text with no semantic marker. Multi-column layouts interleave when parsed linearly. Tables become runs of loose numbers. Charts are images with no alternative text. A retrieval system can often get something out of a PDF, but it gets less, less reliably, and with a worse sense of what the document is actually claiming. When the same content exists as HTML somewhere else on the web, the HTML version wins the citation on structure alone.
The crawler side compounds it. Fetch behavior differs meaningfully across AI user agents, and our work on what AI crawlers actually fetch found that plenty of them are far more selective than Googlebot about non-HTML resources. Bandwidth is not free for them either. A large binary that requires a parsing step is exactly the kind of resource a retrieval pipeline deprioritizes when it has an HTML alternative. So a company that ships its best research as a gated PDF is often absent from AI answers on its own topic while a competitor's thinner blog post gets named.
Which of your PDFs are actually at risk
Not every PDF is a problem, and a blanket migration is a waste of a quarter. Sort your library into four buckets and only two of them need work.
The triage question that sorts most files in about five seconds: if a person read only this document and never visited your site, did you win or lose? Whitepapers and buyer guides lose, because the whole point is to pull the reader into your funnel and a PDF is a dead end with no links. A signed compliance attestation wins, because being a faithful, unchangeable record is exactly its job.
The PDF SEO migration that takes a week
This is a smaller project than it sounds, and it is one of the highest-return weeks a technical team can spend. Pull every PDF URL out of Search Console with impressions or clicks in the last twelve months, and every PDF with referring domains out of your backlink tool. That intersection is your real library. For most mid-market sites it is between 20 and 60 files, not the 400 sitting in the CMS.
For each file in buckets one and two, publish the content as a real HTML page with proper headings, a table of contents, schema, an author byline, and internal links out to the relevant service and product pages. Keep the PDF alive at its original URL so existing links do not break, and add a canonical or a redirect depending on whether the file still needs to be downloadable. Then link the HTML page from the PDF's own first page, so the assets already circulating in inboxes and Slack channels route people back to something indexable.
Two details teams get wrong. First, do not delete the PDFs. Deleting a file with 40 referring domains to fix an indexing concern trades a small problem for a large one. Second, do not make the HTML page a summary that teases the PDF. If the HTML version withholds the substance, engines will not cite it and readers will not link it, and you have rebuilt the same dead end with extra steps. Publish the whole thing.
Do this before your next quarter closes
Open Search Console, filter Page for .pdf, and set the date range to the last 28 days compared against the previous 28. If your PDF impressions fell off a cliff around August 12 or August 18, you are in the affected group and you now have a documented, dated business case for the migration budget you have been quietly asking for. If they did not fall, you got a free warning and a window to act in without an emergency attached.
Then pick the three PDFs with the most referring domains and publish HTML versions of them this month. Not all 40. Three. That gives you a before-and-after on the same domain with the same authority, and it is a far more persuasive internal argument than any external case study, including ours. The teams that ran this migration in 2024 for accessibility reasons are the ones sitting comfortably this week, which is usually how these things go.
If the library is large enough that triage alone is a project, that is what a technical SEO program or a full content marketing engagement is for, and it is a recurring finding for fintech and other regulated categories where half the useful content on the site is legally required to exist as a document. The point is not that PDFs are bad. The point is that a PDF should be a copy of your content, never the only version of it. Google spent two weeks in August demonstrating why, and it did not owe anyone an explanation first.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.