Every GEO conversation eventually gets to the same question, usually a few weeks into an engagement. A team has a genuinely good page, well-researched, accurate, clearly written, ranking respectably. It never shows up in an AI answer. Robots.txt is clean. Schema validates. Nothing's blocked. So what, exactly, is it?
The question everyone asks eventually
The instinctive first move is almost always technical: recheck the crawler access, audit the markup, look for a rendering issue that's hiding the content from a bot. That instinct isn't unreasonable. It's just misallocated, according to the first real attempt anyone has published to actually taxonomize why citations fail in the first place.
A March 2026 paper, built around a framework called AgentGEO, set out to classify citation failures systematically rather than anecdotally, across a real dataset of cases where a topically relevant page should have been citable and wasn't. What it found reorders the whole troubleshooting checklist most GEO audits still run in the order technical-first.
That's not a small correction. It's a full taxonomy, the first attempt anyone's published to actually categorize why a topically relevant page fails to get cited, broken into four buckets that span every stage of what has to go right for a citation to happen: whether a page can be reached at all, whether what's on it is worth citing, whether it actually answers the right version of the question, and whether something else beat it to the same fact. Most GEO troubleshooting checks the first bucket exhaustively and treats the other three as an afterthought, if it gets to them at all.
The taxonomy
| FAILURE CATEGORY | SHARE OF FAILURES | WHAT IT ACTUALLY MEANS |
|---|---|---|
| Semantic alignment | 62.2% | The content answers the wrong version of the question |
| Content quality | 27.1% | The content is too thin or too poorly structured to cite |
| Technical integrity | 10.1% | The page is genuinely unreachable or unparseable |
| Systemic exclusion | 0.6% | A higher-authority source already owns the exact fact |
Technical integrity is the category everyone checks first, and it's also the smallest one by a wide margin: one in ten failures, covering things like firewall-blocked crawlers, content that only renders client-side, corrupted parsing, or pages so cluttered with ads and navigation chrome that the actual content drowns in noise. Real problems, worth fixing, and not remotely where most of the failures actually live.
It's worth sitting with why this category gets checked first anyway, because the reason isn't that anyone thinks it's the biggest problem. It's that it's the easiest one to check. Robots.txt is a single file. Schema validates with a free tool in thirty seconds. Whether a page's core claim actually matches the specific intent of a query cluster requires editorial judgment, real familiarity with what buyers are actually asking, and no automated tool that spits out a pass or fail. Teams check the easy thing first not because it's the likely culprit, but because it's the fast one, and that habit is exactly what this taxonomy argues against.
The category that's actually winning, by a wide margin
Semantic alignment is the one doing the damage, at 62.2% of all failures, more than the other three categories combined. It breaks down into four specific ways a page can miss: intent divergence, where the content answers a transactional question with an informational page or vice versa; contextual gap, where the content is topically related but missing the specific entities or terminology the query actually needs; outdated information, content that's factually stale relative to what the query needs answered right now; and localization mismatch, content built for the wrong region or language entirely.
This reframes a huge share of "why isn't this getting cited" conversations. A page can be perfectly accessible, perfectly well-written on its own topic, and still fail because it's answering an adjacent question to the one the engine's query, or one of its fan-out sub-queries, actually needed answered. That's not a bug in the page. It's a mismatch between what the page was built to say and what a specific query cluster is asking for, and no amount of schema markup or crawler access fixes a mismatch at that level.
Take intent divergence specifically, since it's genuinely the easiest of the four to picture clearly. A page built to explain a category, "what is a CDP," written in the informational register that kind of query rewards, will keep losing citations to a comparison or vendor-shortlist page every time the actual query cluster shifts toward a transactional intent, "which CDP for a mid-market SaaS company." Both pages might rank reasonably for their own respective keywords. Only one of them is the right shape for what a buyer at that specific stage is actually asking, and the wrong-shaped page can be flawless on every technical and structural dimension and still never get picked.
Outdated information works the same underlying way, just approached from a slightly different angle. A page can be perfectly aligned with the right intent and still lose the citation slot to a fresher source the moment the facts inside it drift out of date, a pricing page that hasn't reflected a plan change, a benchmark citing last year's numbers when a newer study exists. None of that is a crawler problem or a structure problem. It's a maintenance problem, and it's invisible to any technical audit that only checks whether a page is reachable, not whether what it says is still true.
Content quality sits second at 27.1%, split between information scarcity, text too shallow to contain an actual citeable fact, and unstructured layout, the content is there but buried in dense prose instead of the kind of scannable structure a model can lift cleanly. That's closer to what GEO advice already targets: structure, extractability, the anatomy of a citation most GEO playbooks already cover. It's real, it's the second-biggest bucket, and it's still smaller than the semantic problem sitting above it.
Information scarcity specifically is worth distinguishing from thinness in the traditional SEO sense. A page can run well past a word-count minimum and still be scarce in this taxonomy's sense if none of those words resolve into a discrete, quotable fact. A thousand words of context, caveats, and background around a topic, with no single sentence a model could lift as a self-contained answer, scores as scarce here even though it would pass most word-count-based content-quality checks without issue. The measure isn't length. It's whether there's an actual answer buried somewhere inside the length.
Systemic exclusion, competitive redundancy where a higher-authority source already covers the identical fact, or the relevant passage simply gets truncated outside a model's context window, rounds out the taxonomy at a genuinely tiny 0.6%. That's a useful number on its own: the fear that a bigger competitor is simply crowding smaller sites out of citations by default is, per this data, a minor factor next to the other three.
That number deserves a second look precisely because it cuts against a common, comfortable excuse. "We can't compete with a bigger domain on this topic" is an easy story to tell when a citation doesn't land, and it's occasionally true. But at 0.6% of classified failures, it's a rare explanation, not a default one. Most of the time a page loses a citation slot not because a bigger competitor already owns the fact, but because the page itself has a semantic, quality, or access problem that has nothing to do with domain size at all. That's actually better news than it sounds: a domain-size problem is largely unfixable in the short term. A semantic-alignment problem is fixable by next quarter.
What to actually fix, and in what order
Run the audit in the order the data actually supports, not the order instinct suggests, and not the order that just happens to be fastest to check off a list. Start with semantic alignment, because it's the majority of the problem: for your highest-value pages, ask honestly whether the content answers the exact question a buyer would ask, in the terminology they'd actually use, with information that's current, not whether it's technically reachable. A page written to explain a category in general terms, when the actual query cluster wants a specific comparison or a specific number, will keep losing citations no matter how clean its markup is.
Move to content quality second: is there an actual citeable fact in the piece, stated plainly enough to lift on its own, or is the real information buried three qualifying clauses deep in a long, hedged paragraph? This is the level structural optimization work and topical depth both operate on, and it's genuinely worth the effort, just not the first stop.
Check technical integrity last, not because it doesn't matter, but because it's the smallest lever and the easiest one to verify quickly once the bigger problems are addressed. A GEO audit that starts and ends with robots.txt and schema validation is auditing roughly a tenth of the actual problem space and calling the job done, when the other nine-tenths of it never got looked at in the first place.
The framework's headline result backs up why this ordering matters: AgentGEO, built around this exact taxonomy, achieved a 40%-plus relative improvement in citation rate while modifying only 5% of a page's content, against a 25% modification rate for baseline optimization approaches that weren't diagnosing which failure mode they were actually looking at. In its published benchmark, AgentGEO reached a 79.52% citation rate versus 68.80% for the strongest baseline it was tested against, a real, meaningful, double-digit percentage-point gap. Knowing which of the four categories a page is actually failing on, before touching anything, is what makes that kind of efficiency possible instead of guessing your way through a rewrite and hoping something sticks.
That efficiency gap is the real argument for diagnosing before editing. A baseline approach that doesn't classify the failure first ends up rewriting five times more of the page than it needs to, because it's guessing at what might help rather than targeting the actual cause. A page failing on outdated information doesn't need its structure rebuilt. A page failing on unstructured layout doesn't need new facts added. Treating every citation failure as a generic "make it better" problem is exactly the instinct this taxonomy exists to correct, and the 5%-versus-25% content-modification gap is what that correction is worth in practice, not just in theory.
The full paper, with the complete breakdown of subcategories and the methodology behind how the taxonomy was built and validated, is published on arXiv for anyone who wants the underlying detail rather than a summary of the headline percentages: Diagnosing and Repairing Citation Failures in Generative Engine Optimization.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.