Something Inc.LoginSchedule a free consultation
RESEARCH

How AI engines actually decide who to cite: a four-stage model

Five independently published 2026 studies each answer a different piece of the citation question. Stacked in order, they describe an actual decision pipeline, not a single ranking factor, and most GEO troubleshooting starts at the wrong stage of it.

RESEARCH5 SOURCES SYNTHESIZED
36.6%
of Claude prompts trigger live web search at all (Profound)
0.61%
of retrieved candidates OpenAI actually selects as Reddit sources (DEJAN)
62.2%
of citation failures trace to semantic misalignment, not access (AgentGEO)
5x
citation-rate lift once a brand is already the chosen recommendation (Seer)
METHODOLOGYSomething Inc. did not run a new primary study for this piece. We cross-referenced five independently published 2026 studies, each measuring a different stage of how an AI engine goes from receiving a query to citing a source: Josh Blyskal/Profound on live-search frequency, Dan Petrovic/DEJAN on per-engine source-selection behavior, the AgentGEO academic taxonomy of citation failure modes, Princeton's GEO tactic-lift framework via ConvertMate's 2026 benchmark, and Seer Interactive's ghost-citation study on recommendation-versus-citation ordering. Each source retains its own original methodology, sample, and time window; we did not re-run or independently verify their underlying data collection. What follows sequences these five findings into a single four-stage model of the citation decision, in the order the underlying mechanics actually occur, and explains why most GEO troubleshooting checks the stages out of order.

Most GEO advice treats "getting cited" as one decision an AI engine makes. It's closer to four separate decisions, made by different mechanisms, often in different systems entirely, and a page can pass three of the four and still never get cited because it failed the one nobody checked first.

Methodology

The five studies behind this model were selected because each isolates a genuinely different stage of the citation pipeline rather than re-measuring the same thing from a new angle. Profound's data answers whether an engine searches the live web at all for a given prompt. DEJAN's data answers, given that it does search, which sources it actually selects from what it retrieves. The AgentGEO taxonomy answers why a specific piece of content, once retrieved and considered, fails to get cited. Princeton's framework, and the ConvertMate benchmark corroborating it, measures what a specific class of content-level edit is worth once a page has cleared the earlier stages. And Seer's ghost-citation research answers a question that sits partly outside the other four: whether the citation decision was ever really contingent on retrieval at all, or whether the underlying recommendation was already set before retrieval started. Stacking these five in sequence, rather than treating each as a stand-alone finding, is the actual contribution of this piece; the underlying data belongs to five different research teams.

The first branch point in the pipeline is the one most GEO programs skip entirely: does the engine consult the live web for this query at all, or answer purely from what it already learned in training. Josh Blyskal's Profound research, published mid-2026, found Claude invokes live web search for only 36.6% of tested prompts. For the majority of queries, Claude answers from trained or cached knowledge with no live retrieval step whatsoever, which means no citation opportunity exists for that query at all, regardless of how well any page is optimized.

This single number reframes a huge share of GEO troubleshooting. A team investigating why a well-structured, well-sourced page isn't getting cited by Claude on a specific query needs to first rule out the possibility that Claude never searched for that query in the first place. No amount of on-page optimization moves a citation decision that a live-retrieval step never reached. This is also the stage where engines diverge from each other most sharply and most consequentially: a page's Claude strategy and its ChatGPT strategy may need to start from entirely different assumptions about whether retrieval happens at all, not just how it's weighted once it does.

It's worth being precise about what "answers from trained knowledge" actually means for a brand trying to get cited, because it's easy to misread as "doesn't matter." A query answered from parametric memory without live retrieval isn't a query where content optimization is irrelevant, it's a query where the relevant content already happened, months or years earlier, whenever the material that shaped the model's training data was originally published and indexed somewhere the training pipeline could reach it. That reframes stage 1 from "citation opportunity: none" to "citation opportunity: already closed, or already won, before this specific query was ever run." The practical difference matters for where a team spends effort: chasing a citation on a query type that skips retrieval entirely is a wasted cycle; building the kind of durable, widely-corroborated presence that shapes future training data is not, even though neither shows up as a citation on today's tracked query.

This also means the 36.6% figure isn't a ceiling on how much of Claude's total output cites external sources, it's a ceiling on how much of it can be influenced by anything published or updated after the model's last training cutoff, on a per-query basis. A brand's content strategy for an engine with a low live-search rate has to account for both timelines at once: the live-retrieval share, where fresh content and technical access still matter in something close to the way classic GEO advice assumes, and the parametric share, where the more relevant lever is what got baked into training months ago and won't be affected by anything published today.

Stage 2: Which sources does it trust

Once an engine does search live, the next branch point is which sources it actually selects from what retrieval surfaces, and Dan Petrovic's DEJAN research, covered here previously, found this varies by engine more sharply than most unified GEO checklists account for. Across 139,601 sources analyzed, OpenAI selected Reddit as a source in just 0.61% of retrieved candidates, while Google's selection rate for the same source type ranged from 9% to 60% depending on query type, and Anthropic's rate was effectively 0%. Gemini showed a different pattern entirely, selecting whichever page it read first in 92% of cases, a recency-of-processing effect rather than a trust-ranking one.

ENGINESOURCE-SELECTION BEHAVIORWHAT IT IMPLIES FOR STRATEGY
OpenAI / ChatGPTSelects Reddit in 0.61% of retrieved candidatesCommunity-platform citations are a weak lever for this engine specifically
Google (AI Mode / Overviews)9-60% selection rate for the same source type, by queryTrust varies heavily by query type: no single blended assumption holds
Anthropic / Claude~0% selection rate for the same source typeThe engine that searches least also filters hardest once it does
GeminiCites whichever page it reads first, 92% of the timeProcessing order, not authority ranking, is the dominant factor

The strategic implication is blunt: a single, blended source-trust assumption, "community platforms help GEO" or "community platforms don't matter," is wrong for at least two of these four engines no matter which way you round it. Stage 2 has to be audited per engine, not once for "AI search" as a category, and the source types worth investing in differ by which specific engine a program is actually trying to win.

Gemini's pattern deserves its own note, because it's mechanically different from the other three rather than just quantitatively different. OpenAI, Google's broader AI Mode, and Anthropic all appear to be running something closer to a trust-weighted selection process, different weights, but the same basic shape, where a source type's historical reliability shapes how often it gets chosen. Gemini's 92% first-read-wins pattern describes something closer to a processing-order effect than a trust ranking at all. That distinction matters operationally: a source-authority campaign aimed at improving Gemini citation share is optimizing for the wrong variable if the actual lever is retrieval order rather than perceived trustworthiness. Speed and structural accessibility, how quickly and cleanly a page can be parsed once retrieved, may matter more for this specific engine than the credibility signals that move the other three.

Stage 3: Does the content fit

A source that clears stages 1 and 2, the engine searched live, and the source type is one that engine trusts, still has to pass a content-level fit test, and this is where the AgentGEO taxonomy, covered in detail on this site, does the most useful work. Sampling failures directly rather than testing interventions on hand-picked content, AgentGEO found 62.2% of citation failures trace to semantic alignment problems, intent divergence, contextual gaps, outdated information, localization mismatch, far ahead of content quality issues at 27.1% and technical access problems at 10.1%.

Semantic alignment62%
Content quality27%
Technical integrity10%
Systemic exclusion1%

Share of citation failures by category (AgentGEO taxonomy)

This is also the stage where the Princeton GEO framework's tactic-level findings, statistics addition and cited sources each adding roughly 30% visibility lift, which Something Inc. examined critically elsewhere this week, actually operate. Those tactics live inside the content-quality slice of stage 3, a real but comparatively small piece of the overall failure distribution. A content team optimizing only at this level, adding statistics and citations without first confirming the content semantically matches the query, is polishing the smaller of the two content-level failure categories while leaving the larger one, alignment, untouched.

Stage 4: Was the brand already picked

The fourth stage is the one that complicates the tidy story the first three stages tell, because it suggests the citation decision may, for a meaningful share of queries, have already been effectively made before stages 1 through 3 ever ran. Seer Interactive's ghost-citation study, published in March 2026 from 362,188 analyzed responses, found a brand's citation rate jumps to 53.1% once that same response already recommends it, versus 10.6% when it doesn't, a 5x gap running in the direction of recommendation causing citation rather than the reverse.

Read against the first three stages, this suggests parametric memory, what the model already "knows" about a brand from training, can function as something like a pre-filter sitting upstream of live retrieval entirely. A brand well-represented in a model's training data may get recommended by default, with retrieval mostly serving to find supporting citations for a decision the parametric layer already made. A brand poorly represented in training data may clear stages 1 through 3 on a specific page, live search happened, the source type is trusted, the content fits, and still not get recommended, because the stage 4 pre-selection never favored it in the first place.

This is the stage most GEO strategy has no real tooling for, because it isn't addressable through anything published today. It's shaped by the aggregate signal a brand accumulated across the entire window of content the model was trained on, third-party mentions, review volume, press coverage, community discussion, all compounded over time rather than earned through any single campaign. That's an uncomfortable answer for a discipline built around measurable, campaign-length interventions, but it's consistent with what our own review-profile research already found: 99.5% of review-driven citations arrived because a review page ranked well in ordinary organic search and got pulled into an AI answer from there, meaning the underlying mechanism was durable, indirect authority-building rather than an AI-search-specific tactic at all.

HOW A QUESTION BECOMES A CITATION
Stage 1: Live searchOnly 36.6% of Claude prompts trigger retrieval at all (Profound)
Stage 2: Source trustWhich source types the engine selects from, varies sharply by engine (DEJAN)
Stage 3: Content fit62.2% of failures here are semantic, not technical (AgentGEO)
Stage 4: Parametric pick5x citation lift once the brand is already the chosen answer (Seer)

Why the stage order matters

Sequenced this way, the four stages explain a pattern GEO teams keep running into without a name for it: a page that looks correct by every checkable standard, live-retrievable, from a trusted source type, well-structured and on-topic, that still doesn't get cited consistently. Checked stage by stage, the honest diagnostic order runs opposite to how most technical audits actually proceed. Most teams start at stage 3's easiest-to-verify layer, technical access and on-page structure, because it's mechanically checkable in an afternoon. That's the smallest failure category in the whole model, 10.1% within stage 3 alone. Stage 1 and stage 4 are both upstream of anything a page-level audit can fix at all, and neither shows up in a standard technical crawl.

THE DIAGNOSTIC ORDER THIS MODEL IMPLIESBefore auditing a single page: confirm the engine searches live for the target query class at all (stage 1). Confirm the source type is one that engine trusts for that query (stage 2). Only then audit semantic fit and content quality (stage 3). Treat parametric pre-selection (stage 4) as a separate, slower-moving lever, built through sustained third-party corroboration, not a page-level fix.

Applying the model

For a practical audit, the model reorders where GEO budget should go relative to how most programs currently allocate it. Stage 1 argues for query-class research before content production: know, per engine, which of your target queries even trigger live retrieval before building content optimized for a retrieval step that may not happen. Stage 2 argues for per-engine strategy rather than a single blended GEO checklist, since source trust genuinely doesn't generalize across engines. Stage 3 argues for auditing semantic alignment before content-quality polish, the sequencing already laid out in detail elsewhere on this site. Stage 4 argues for treating third-party corroboration and category authority as the slowest-moving, highest-ceiling lever in the model, since it appears to shape outcomes before the other three stages even run.

The four stages also explain why a single citation-rate number, tracked without knowing which stage moved it, is close to useless for diagnosing what to fix next. A citation-rate drop could mean an engine stopped live-searching a query class (stage 1), changed its source-trust weighting (stage 2), or reflect a genuine content problem (stage 3), and distinguishing between them requires instrumentation most reporting and analytics stacks weren't built to capture, since they were designed around a single blended AI-visibility score rather than a staged model like this one.

It's worth stating plainly what this model does and doesn't claim. It doesn't claim these four stages are the only mechanics involved, real systems are messier than any four-box diagram, and it doesn't claim the five source studies used identical methodologies or time windows, they didn't. What it claims is narrower and more useful: that sequencing five real, independently verified findings in the order their underlying mechanics actually run produces a more actionable diagnostic than treating any one of them as the whole story, which is how most of them currently circulate in isolation.

None of these five studies was designed to talk to the others. Profound measured retrieval frequency. DEJAN measured source selection. AgentGEO measured failure taxonomy. Princeton and ConvertMate measured tactic-level lift. Seer measured recommendation-citation ordering. Put in sequence, they describe a single pipeline a query actually travels through, and a GEO program built around auditing that pipeline in order, rather than starting wherever a checklist happens to start, is the version of this research worth acting on. For a B2B SaaS team running its next quarterly GEO audit, the four-stage model is a more useful starting checklist than any single one of the five studies behind it, because it's the only one of the five that tells you which stage to check first.

Cite this research● LIVE
Josh Blyskal - Profound - "The state of AEO in 2026: Claude is not ChatGPT" - Jul 22, 2026
https://www.tryprofound.com/resources
Dan Petrovic - DEJAN - source-selection analysis across 139,601 sources - Jul 21 & 26, 2026
https://dejan.ai/
AgentGEO - "A Taxonomy of Citation Failure in Generative Engines" - arXiv 2603.09296 - Mar 10, 2026
https://arxiv.org/abs/2603.09296
ConvertMate - "GEO Benchmark Study 2026: What Actually Drives Visibility in Generative Search" - 2026
https://www.convertmate.io/research/geo-benchmark-2026
John Lovett - Seer Interactive - "LLM Ghost Citations: Why Your Content Is Working and Your Brand Isn't" - Mar 24, 2026
https://www.seerinteractive.com/insights/llm-ghost-citations-why-your-content-is-working-and-your-brand-isnt

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

TT
Tyler TruffiMANAGING PARTNER, SOMETHING INC.

Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.