Most GEO advice treats "getting cited" as one decision an AI engine makes. It's closer to four separate decisions, made by different mechanisms, often in different systems entirely, and a page can pass three of the four and still never get cited because it failed the one nobody checked first.
Methodology
The five studies behind this model were selected because each isolates a genuinely different stage of the citation pipeline rather than re-measuring the same thing from a new angle. Profound's data answers whether an engine searches the live web at all for a given prompt. DEJAN's data answers, given that it does search, which sources it actually selects from what it retrieves. The AgentGEO taxonomy answers why a specific piece of content, once retrieved and considered, fails to get cited. Princeton's framework, and the ConvertMate benchmark corroborating it, measures what a specific class of content-level edit is worth once a page has cleared the earlier stages. And Seer's ghost-citation research answers a question that sits partly outside the other four: whether the citation decision was ever really contingent on retrieval at all, or whether the underlying recommendation was already set before retrieval started. Stacking these five in sequence, rather than treating each as a stand-alone finding, is the actual contribution of this piece; the underlying data belongs to five different research teams.
Stage 1: Does the engine even search
The first branch point in the pipeline is the one most GEO programs skip entirely: does the engine consult the live web for this query at all, or answer purely from what it already learned in training. Josh Blyskal's Profound research, published mid-2026, found Claude invokes live web search for only 36.6% of tested prompts. For the majority of queries, Claude answers from trained or cached knowledge with no live retrieval step whatsoever, which means no citation opportunity exists for that query at all, regardless of how well any page is optimized.
This single number reframes a huge share of GEO troubleshooting. A team investigating why a well-structured, well-sourced page isn't getting cited by Claude on a specific query needs to first rule out the possibility that Claude never searched for that query in the first place. No amount of on-page optimization moves a citation decision that a live-retrieval step never reached. This is also the stage where engines diverge from each other most sharply and most consequentially: a page's Claude strategy and its ChatGPT strategy may need to start from entirely different assumptions about whether retrieval happens at all, not just how it's weighted once it does.
It's worth being precise about what "answers from trained knowledge" actually means for a brand trying to get cited, because it's easy to misread as "doesn't matter." A query answered from parametric memory without live retrieval isn't a query where content optimization is irrelevant, it's a query where the relevant content already happened, months or years earlier, whenever the material that shaped the model's training data was originally published and indexed somewhere the training pipeline could reach it. That reframes stage 1 from "citation opportunity: none" to "citation opportunity: already closed, or already won, before this specific query was ever run." The practical difference matters for where a team spends effort: chasing a citation on a query type that skips retrieval entirely is a wasted cycle; building the kind of durable, widely-corroborated presence that shapes future training data is not, even though neither shows up as a citation on today's tracked query.
This also means the 36.6% figure isn't a ceiling on how much of Claude's total output cites external sources, it's a ceiling on how much of it can be influenced by anything published or updated after the model's last training cutoff, on a per-query basis. A brand's content strategy for an engine with a low live-search rate has to account for both timelines at once: the live-retrieval share, where fresh content and technical access still matter in something close to the way classic GEO advice assumes, and the parametric share, where the more relevant lever is what got baked into training months ago and won't be affected by anything published today.
Stage 2: Which sources does it trust
Once an engine does search live, the next branch point is which sources it actually selects from what retrieval surfaces, and Dan Petrovic's DEJAN research, covered here previously, found this varies by engine more sharply than most unified GEO checklists account for. Across 139,601 sources analyzed, OpenAI selected Reddit as a source in just 0.61% of retrieved candidates, while Google's selection rate for the same source type ranged from 9% to 60% depending on query type, and Anthropic's rate was effectively 0%. Gemini showed a different pattern entirely, selecting whichever page it read first in 92% of cases, a recency-of-processing effect rather than a trust-ranking one.
| ENGINE | SOURCE-SELECTION BEHAVIOR | WHAT IT IMPLIES FOR STRATEGY |
|---|---|---|
| OpenAI / ChatGPT | Selects Reddit in 0.61% of retrieved candidates | Community-platform citations are a weak lever for this engine specifically |
| Google (AI Mode / Overviews) | 9-60% selection rate for the same source type, by query | Trust varies heavily by query type: no single blended assumption holds |
| Anthropic / Claude | ~0% selection rate for the same source type | The engine that searches least also filters hardest once it does |
| Gemini | Cites whichever page it reads first, 92% of the time | Processing order, not authority ranking, is the dominant factor |
The strategic implication is blunt: a single, blended source-trust assumption, "community platforms help GEO" or "community platforms don't matter," is wrong for at least two of these four engines no matter which way you round it. Stage 2 has to be audited per engine, not once for "AI search" as a category, and the source types worth investing in differ by which specific engine a program is actually trying to win.
Gemini's pattern deserves its own note, because it's mechanically different from the other three rather than just quantitatively different. OpenAI, Google's broader AI Mode, and Anthropic all appear to be running something closer to a trust-weighted selection process, different weights, but the same basic shape, where a source type's historical reliability shapes how often it gets chosen. Gemini's 92% first-read-wins pattern describes something closer to a processing-order effect than a trust ranking at all. That distinction matters operationally: a source-authority campaign aimed at improving Gemini citation share is optimizing for the wrong variable if the actual lever is retrieval order rather than perceived trustworthiness. Speed and structural accessibility, how quickly and cleanly a page can be parsed once retrieved, may matter more for this specific engine than the credibility signals that move the other three.
Stage 3: Does the content fit
A source that clears stages 1 and 2, the engine searched live, and the source type is one that engine trusts, still has to pass a content-level fit test, and this is where the AgentGEO taxonomy, covered in detail on this site, does the most useful work. Sampling failures directly rather than testing interventions on hand-picked content, AgentGEO found 62.2% of citation failures trace to semantic alignment problems, intent divergence, contextual gaps, outdated information, localization mismatch, far ahead of content quality issues at 27.1% and technical access problems at 10.1%.
Share of citation failures by category (AgentGEO taxonomy)
This is also the stage where the Princeton GEO framework's tactic-level findings, statistics addition and cited sources each adding roughly 30% visibility lift, which Something Inc. examined critically elsewhere this week, actually operate. Those tactics live inside the content-quality slice of stage 3, a real but comparatively small piece of the overall failure distribution. A content team optimizing only at this level, adding statistics and citations without first confirming the content semantically matches the query, is polishing the smaller of the two content-level failure categories while leaving the larger one, alignment, untouched.
Stage 4: Was the brand already picked
The fourth stage is the one that complicates the tidy story the first three stages tell, because it suggests the citation decision may, for a meaningful share of queries, have already been effectively made before stages 1 through 3 ever ran. Seer Interactive's ghost-citation study, published in March 2026 from 362,188 analyzed responses, found a brand's citation rate jumps to 53.1% once that same response already recommends it, versus 10.6% when it doesn't, a 5x gap running in the direction of recommendation causing citation rather than the reverse.
Read against the first three stages, this suggests parametric memory, what the model already "knows" about a brand from training, can function as something like a pre-filter sitting upstream of live retrieval entirely. A brand well-represented in a model's training data may get recommended by default, with retrieval mostly serving to find supporting citations for a decision the parametric layer already made. A brand poorly represented in training data may clear stages 1 through 3 on a specific page, live search happened, the source type is trusted, the content fits, and still not get recommended, because the stage 4 pre-selection never favored it in the first place.
This is the stage most GEO strategy has no real tooling for, because it isn't addressable through anything published today. It's shaped by the aggregate signal a brand accumulated across the entire window of content the model was trained on, third-party mentions, review volume, press coverage, community discussion, all compounded over time rather than earned through any single campaign. That's an uncomfortable answer for a discipline built around measurable, campaign-length interventions, but it's consistent with what our own review-profile research already found: 99.5% of review-driven citations arrived because a review page ranked well in ordinary organic search and got pulled into an AI answer from there, meaning the underlying mechanism was durable, indirect authority-building rather than an AI-search-specific tactic at all.
Why the stage order matters
Sequenced this way, the four stages explain a pattern GEO teams keep running into without a name for it: a page that looks correct by every checkable standard, live-retrievable, from a trusted source type, well-structured and on-topic, that still doesn't get cited consistently. Checked stage by stage, the honest diagnostic order runs opposite to how most technical audits actually proceed. Most teams start at stage 3's easiest-to-verify layer, technical access and on-page structure, because it's mechanically checkable in an afternoon. That's the smallest failure category in the whole model, 10.1% within stage 3 alone. Stage 1 and stage 4 are both upstream of anything a page-level audit can fix at all, and neither shows up in a standard technical crawl.
Applying the model
For a practical audit, the model reorders where GEO budget should go relative to how most programs currently allocate it. Stage 1 argues for query-class research before content production: know, per engine, which of your target queries even trigger live retrieval before building content optimized for a retrieval step that may not happen. Stage 2 argues for per-engine strategy rather than a single blended GEO checklist, since source trust genuinely doesn't generalize across engines. Stage 3 argues for auditing semantic alignment before content-quality polish, the sequencing already laid out in detail elsewhere on this site. Stage 4 argues for treating third-party corroboration and category authority as the slowest-moving, highest-ceiling lever in the model, since it appears to shape outcomes before the other three stages even run.
The four stages also explain why a single citation-rate number, tracked without knowing which stage moved it, is close to useless for diagnosing what to fix next. A citation-rate drop could mean an engine stopped live-searching a query class (stage 1), changed its source-trust weighting (stage 2), or reflect a genuine content problem (stage 3), and distinguishing between them requires instrumentation most reporting and analytics stacks weren't built to capture, since they were designed around a single blended AI-visibility score rather than a staged model like this one.
It's worth stating plainly what this model does and doesn't claim. It doesn't claim these four stages are the only mechanics involved, real systems are messier than any four-box diagram, and it doesn't claim the five source studies used identical methodologies or time windows, they didn't. What it claims is narrower and more useful: that sequencing five real, independently verified findings in the order their underlying mechanics actually run produces a more actionable diagnostic than treating any one of them as the whole story, which is how most of them currently circulate in isolation.
None of these five studies was designed to talk to the others. Profound measured retrieval frequency. DEJAN measured source selection. AgentGEO measured failure taxonomy. Princeton and ConvertMate measured tactic-level lift. Seer measured recommendation-citation ordering. Put in sequence, they describe a single pipeline a query actually travels through, and a GEO program built around auditing that pipeline in order, rather than starting wherever a checklist happens to start, is the version of this research worth acting on. For a B2B SaaS team running its next quarterly GEO audit, the four-stage model is a more useful starting checklist than any single one of the five studies behind it, because it's the only one of the five that tells you which stage to check first.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.