Something Inc.LoginSchedule a free consultation
TECHNICAL SEO

Programmatic SEO risk is site-level, not page-level

Every programmatic build is pitched as a reversible experiment. Google describes the downside as a judgment about your domain, which is the one thing a bulk delete cannot undo.

TECHNICAL SEOCONTRARIANSCALED CONTENT

Programmatic SEO gets approved on the strength of its exit. Spin up forty thousand pages from a template and a database, watch the impressions for a quarter, and if it does not work, delete the directory. That framing is why programmatic SEO risk keeps getting underwritten as a page-level bet with a clean undo. Google has just described the downside in different terms, and the difference is the whole argument: the thing at stake is not the page set. It is the domain the page set is attached to.

The specific exchange happened on Bluesky in early September, when someone described a site generating pages from combinations of domain names, technologies and attributes and asked why it was not performing. John Mueller's answer, reported by Search Engine Journal on September 8 and by Search Engine Roundtable the day before, contained one sentence that should change how these projects get budgeted. His systems, he said, have possibly lost faith in your site providing good value to users based on the old pages.

22%
of roughly 400 sites Glenn Gabe tracked after the September 2023 helpful content update had regained 20 percent or more of their lost traffic by August 2024
129 of 130
of the hardest hit sites in Lily Ray's analysis had only lost further visibility since, with a single site up 3 percent
The site
what Mueller named as the thing his systems lost faith in, rather than the generated pages that caused it
Months
the horizon Mueller put on resolving it, by comparison with spam action and core update recoveries

Read those together and the standard approval memo stops working. The memo assumes the asset and the liability live at the same address. They do not. The asset is a directory of URLs. The liability is an assessment of a host, and hosts are not deletable.

The subject of that sentence was the site, not the pages

Grammar is doing real work here, so it is worth slowing down on it. Mueller did not say the pages had lost value, or that the pages would stop ranking, or that the pages were spam. He said the systems had lost faith in the site, based on the old pages. The pages are the evidence in that sentence. The site is the defendant.

He was careful about it too, and honest practitioners should be equally careful repeating it. He said possibly. He named no ranking system, no threshold, no number of pages at which this triggers, and no recovery timeline beyond a comparison to how long spam and core update recoveries take, which he characterised as sometimes many months. Anyone selling you a precise page count as the safe ceiling is inventing it. What is on the record is the direction of the effect and where it lands, and both of those are enough to change a decision.

He also refused the easy generalisation, which matters for anyone whose whole category depends on templated pages. Mueller's phrasing was programmatic SEO like this, aimed at the specific setup in front of him: pages assembled from attribute permutations, each carrying, in his words, some value, with an overall picture that is not that exciting. Location pages for a business with real locations are not that. A specification database with genuinely distinct records is not that. The failure mode he described is permutation without substance, and the tell is that no individual page would survive being asked what it is for.

Deleting a generated page set closes the tab. It does not close the position. The URLs were the evidence, and the evidence is not what gets ranked.

Programmatic SEO risk does not leave when the URLs do

The practical consequence is a mismatch between what a cleanup removes and what a cleanup is supposed to fix. Teams run the delete, watch the index count fall, and read that falling number as recovery in progress. It is not. It is the removal of the evidence, which is necessary and not sufficient. Here is what actually separates when you pull the directory.

WHAT THE BUILD CREATEDWHERE IT LIVESGONE WHEN YOU DELETE THE URLSSTILL IN PLAY AFTERWARDS
The generated page setIndividual URLs and their templateYes, immediatelyNothing. This is the only genuinely clean part of the rollback
Index bloat and crawl wasteGoogle's crawl scheduling for the hostGradually, as URLs drop outThe crawl history that taught the scheduler what this host publishes
Site-level quality assessmentThe host, not any URLNoWhatever conclusion the systems already reached about the domain
Internal links routed into the setYour own architectureYes, along with the pages they fedHub pages and navigation built around a set that no longer exists
External links earned by generated pagesThird party sites you do not controlNo, they become errors or redirectsRedirect chains inheriting the same quality question they started with

Row three is the expensive one and it is the row nobody models. Rows one, two, four and five are cleanup tasks with owners and end dates. Row three is a standing position you cannot reach from your own admin panel, and Google's published guidance says so directly: its scaled content abuse policy is written around purpose and value rather than volume or production method, and its guidance on creating helpful content states that a site carrying relatively high amounts of unhelpful content overall can see the whole site affected. Those two sentences have been public for years. The programmatic pitch deck has been quietly ignoring both.

THE REFRAMEA programmatic build is not a reversible experiment on a subdirectory. It is a bet placed with the credit of the entire domain, on terms you cannot renegotiate after the fact. That is a fine bet to make when the pages carry real, differentiated substance. It is a terrible bet to make because the build was cheap, and cheap is now the only thing code generators have made cheaper. If you are not sure which side of that line a plan sits on, that is exactly the question our SEO and GEO audits exist to answer before the pages ship rather than after.

What site-level recovery actually looks like in the data

Mueller's comparison to spam and core update recovery is the useful part of his answer, because unlike the trust mechanism itself, that comparison has public evidence behind it. The most complete picture still comes from the site-level quality event that ran hardest at scale, the September 2023 helpful content update, and from two practitioners who tracked affected sites for a year afterwards instead of for a news cycle.

Gabe cohort, roughly 400 sites: regained 20 percent or more of lost traffic22%
Gabe cohort: had not regained that much by August 202478%
Ray cohort, 130 hardest hit sites: visibility up1%
Ray cohort: visibility down further99%

Two separately tracked cohorts of sites affected by the September 2023 helpful content update, each shown as a percentage of its own cohort. Glenn Gabe followed roughly 400 heavily affected sites and reported in August 2024 how many had regained 20 percent or more of their lost traffic. Lily Ray analysed the 130 hardest hit sites and found one with positive visibility growth. Different samples, different definitions of recovery, same direction.

Two caveats belong on that chart before anyone quotes it. These are observational cohorts selected for having been hit hard, not random samples of the web, so they describe the tail rather than the average case. And the recovery definitions differ: Gabe's 20 percent threshold is a deliberately generous bar, while Ray's measure is directional visibility. Neither is a controlled study and neither author claimed otherwise.

They still settle the question that matters for a budget. The optimistic case in the better-defined dataset is roughly one site in five clawing back a fifth of what it lost, over a period measured in quarters, by teams who knew exactly what had happened to them and were working on it full time. That is the distribution you are buying into. Price the build against that, not against the version where you delete the directory in March and trade normally by April. We walked through what the correction looks like from the enforcement side in a scaled content manual action case study, and the pattern there was the same: the pages came down fast and the standing took far longer.

Why programmatic SEO risk compounds before anyone notices

The uncomfortable part is that these builds look successful during the window when the exposure is being created. Impressions arrive almost immediately, because a large set of long-tail permutations will match a large set of rare queries whatever their quality. That early signal is read as validation, funds the next expansion, and buys the project another quarter before anyone asks about the denominator.

Impressions arrive before judgment doesA permutation set matches rare queries on day one, which produces a chart that goes up. Site-level quality assessment is slower and much less visible. The gap between those two clocks is where the expansion decision gets made, and it is always made on the fast chart.
The marginal page is always the cheapest thing on the roadmapAdding ten thousand more rows costs a template change. Adding one genuinely useful page costs a person a day. Every incentive in the build points at the permutation count, and code generators have widened that gap rather than closed it.
Ratio, not count, is the thing being assessedGoogle's helpful content guidance talks about how much of a site is unhelpful, not how many pages are. Forty thousand generated pages on a two hundred page site of real substance is not the same exposure as the same forty thousand on a two thousand page one, and no team tracks that ratio as a metric.
The whole category got louder at the same timeEvery competitor read the same posts and shipped the same volume in the same year. That is the saturation dynamic we unpacked in why more output stopped improving performance, and it means the generated set is competing against other generated sets while also pulling down the host that published it.

There is a second-order version of this for anyone hosting third party or partner content on their own domain, which is the same structural bet with a different author. Google's site reputation abuse policy treats that as a site-level matter too, and we went through how it plays out across regions in the site reputation abuse playbook. If you would not put your domain behind a permutation set you generated yourself, you should not put it behind one someone paid you to host.

How to underwrite a generated build before you ship it

None of this makes templated page generation illegitimate. Plenty of the best pages on the web are assembled from structured data, and a real database of genuinely distinct records is one of the strongest assets a site can have. The question is never whether the pages are generated. It is whether each one would survive being looked at by a person who does not work for you.

01Underwrite it at the host, not the directoryWrite the downside case as a site-level event, because that is how Google described it. If the honest sentence is that a failed build could cost the domain its standing for two to four quarters, the project needs the approval that sentence implies, and probably a different owner than whoever owns the traffic target.
02Set the ratio before you set the page countDecide the maximum share of indexable URLs the generated set may occupy, and treat it as a hard cap rather than a guideline. That number protects the denominator Google's own guidance describes, and it is the only lever in the build that anyone can defend in a meeting.
03Apply the single page test to a random sample, not the demoPull twenty rows at random and ask what a visitor gets from each that they could not get from the parent page or from a competitor. Demos are always built from the best rows. If the random sample cannot answer the question, the set is permutation without substance and the count makes it worse rather than better.
04Ship in tranches with a kill criterion written firstRelease a defined slice, then hold. Decide in advance what evidence would stop the next tranche, and make it a quality measure rather than an impression count, because impressions will look fine during exactly the period when the exposure is accumulating.
05Keep the generated set architecturally separableOne directory, its own sitemap, its own internal linking, no navigation dependency. It does not reduce the site-level exposure at all, but it makes the cleanup fast, honest and auditable, and it lets you show a clear before and after when you need to demonstrate what changed.

Where a build is already live and already underperforming, the sequence is unglamorous and the order matters. Establish whether the set is actually the problem before you delete anything, because deletion destroys the evidence you would use to diagnose it. Then cut the set down to the rows with defensible substance rather than removing all of them, fix or consolidate what remains, and only then start the long work of demonstrating value at the site level. That last step is the one Mueller said takes time and significant effort, and it is the one no cleanup script performs. It looks a lot like the correction work we described after the August 2026 spam update: the removals are a week of work and the rebuild is a season of it.

DO THIS NEXTPull your indexable URL count and split it into pages a person wrote and pages a template produced. Put the ratio in front of whoever signs off on the roadmap, then read Google's own scaled content abuse policy and Search Engine Journal's write-up of Mueller's answer against your own build. If the templated share is the larger number and nobody in the room can name what a randomly chosen generated page does for a visitor, you are not running an experiment on a subdirectory. You are carrying a site-level position, and the first useful step is saying so out loud.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

TT
Tyler TruffiMANAGING PARTNER, SOMETHING INC.

Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.