Most GEO reporting treats a citation as the win condition. Get cited enough times and the recommendations should follow, the logic goes, the same way ranking well in Google used to precede getting clicked. A study Seer Interactive published in March 2026 says that causal arrow may run backward for a meaningful share of AI answers, and the gap is large enough to change what a citation dashboard should actually be measuring.
The bibliography, not the brainstorm
John Lovett, VP of Analytics at Seer Interactive, ran six independent behavioral tests across six AI platforms (Claude and Meta were excluded from the citation-specific analysis), collecting 541,213 raw responses and analyzing 362,188 of them in detail, spanning 20 brands, 5 funnel stages, and 8 industry categories over a February 2026 research window. The question behind the study was simple: does getting cited more often predict getting recommended more often, or is it the other way around.
The finding cuts against the tidy version of GEO strategy. When a brand showed up in an AI answer's actual recommendation, that same response cited the brand's own source 53.1% of the time. When the brand wasn't recommended in the answer, its citation rate for that same query dropped to 10.6%. Lovett's phrase for this is worth sitting with: the citations are the bibliography, not the brainstorm. His working hypothesis is that large language models generate a brand recommendation first, pulled from parametric memory, the facts and associations baked into the model during training, and only afterward retrieve supporting sources to back the answer up. Citation, in that sequence, is closer to a footnote than a cause.
Where the gap is worst
Lovett's team also measured what they call the "ghost citation" rate: how often a brand gets cited in an answer that doesn't actually recommend it. That rate isn't uniform. It's highest at the awareness stage of the funnel, at 5.0%, the point in a buyer's journey where an AI engine is most likely to cite a source for background context without committing to naming a specific brand as the answer. It's lowest in categories with tighter, more established competitive sets.
| SEGMENT | GHOST CITATION RATE | WHAT IT MEANS |
|---|---|---|
| Awareness-stage queries | 5.0% | Highest ghost rate in the funnel: cited for context, not recommended |
| Industrial Services | 0.3% | Lowest ghost rate measured: citation and recommendation track tightly |
| Financial Services / HR Tech | Under 2% | Established, well-differentiated categories show less citation/recommendation drift |
| Hospitality & Travel | 20+ point spread | Largest gap between best- and worst-performing brands in the same category |
The pattern across categories suggests ghost citations aren't random noise, they track how settled a category's competitive field is inside the model's training data. In categories like industrial services, where a small number of established players dominate the source material a model would have encountered, the model's recommendation and its citation choice tend to agree with each other. In categories with a wide, fragmented field of near-identical options, like a lot of hospitality and travel content, the model has more room to cite a source without that source actually being the thing it recommends.
The methodology behind that split is worth understanding before applying the numbers to your own reporting. Lovett's team didn't run one large query set and call it done; they ran six separate behavioral tests, each isolating a different variable, funnel stage, industry category, brand tier, and query phrasing among them, then checked whether the citation/recommendation pattern held across all six. It did. That consistency across independently designed tests is what makes the 53.1%-versus-10.6% split worth taking seriously instead of filing it as one report's quirky finding. A single test showing a big gap could be an artifact of how that particular test was built. Six tests agreeing, across six different platforms and 20 different brands, is a much harder result to explain away as measurement noise, and it's worth reading Lovett's original write-up directly for the full breakdown of how each test was isolated.
What this breaks about GEO measurement
This is the same distinction our own anatomy-of-an-ai-citation research already draws between mention rate and citation rank, extended one level further. A blended citation count was already an incomplete metric because it didn't separate how often you appear from where you rank in the source list. Lovett's data adds a third failure mode: it doesn't separate being cited from being recommended at all. A page can be a frequently-cited source on a topic without the brand behind it ever actually being the answer, and a dashboard that only counts citation frequency has no way to tell the difference.
That has a direct, uncomfortable implication for content strategy. If a meaningful share of citations arrive after the recommendation decision is already made rather than causing it, then publishing more citable content on a topic where your brand isn't already a strong candidate in the model's training data may increase your citation count without moving the number that actually drives pipeline. It's the AI-search equivalent of ranking on page one for a term nobody searches with buying intent: technically a win, functionally hollow.
It also changes how a reasonable person should read a month-over-month citation report. A rising citation count looks like unambiguous progress on every dashboard built to track it as a single line. This data says a rising count could mean one of two very different things: either your brand is genuinely winning more recommendations and picking up citations as a natural side effect, or your content is getting swept into more answers as background context without the underlying recommendation ever shifting in your favor. Those are not the same win, and they don't call for the same next move. The first case means keep doing what's working. The second means the content strategy is running ahead of the brand-authority work that would actually make the model pick you.
How to run the co-occurrence audit
The check itself doesn't require new tooling, just a different question asked of data most GEO programs are already collecting. Pull the last 60 to 90 days of tracked queries where your brand appears as a cited source, the same query set feeding your existing citation-rate dashboard. For each one, read the actual generated answer, not just the citation log, and mark whether your brand is named as part of the recommendation itself or only linked as a background source. That single pass turns a blended citation count into two separate numbers: a true recommendation-linked citation rate and a ghost-citation rate.
Segment the results the same way Lovett's team did, by funnel stage and by category, and expect the same shape to show up. Awareness-stage queries and fragmented, undifferentiated categories will carry the highest ghost rates in almost any brand's data, because that's where the model has the least training-data signal pointing it toward one specific answer. Bottom-funnel, comparison-style queries in categories where your brand already has a clear position will tend to show the tighter citation/recommendation coupling closer to the 53.1% figure. If your own numbers invert that pattern, showing high ghost rates on queries where your brand should already be well established, that's a more specific signal than an aggregate citation count would ever surface: something about how the model represents your brand in that category doesn't match how a human buyer would rank you.
“The citations are the bibliography, not the brainstorm.”
What to do about it
The practical fix isn't abandoning citation tracking, it's stopping the treatment of citation count as a stand-alone success metric. Pull a sample of your highest-citation queries and check, response by response, whether your brand is the one actually being recommended or just the one being footnoted. Our failure-taxonomy breakdown of why citations fail found semantic alignment, not technical access, causes the majority of citation misses; this data suggests the reverse failure, citation without recommendation, deserves the same scrutiny rather than being counted as success by default.
Third-party signal strength is the more durable lever here, since it shapes the training-data associations the model draws its initial pick from long before any single page gets crawled for a specific query. Our review-profile citation study already showed how heavily models weight third-party validation, review volume in that case, when deciding what to surface; that's evidence for the same mechanism Lovett's data points to, a brand's standing outside its own content shaping whether it gets picked at all. Building GEO programs and broader content strategy around third-party corroboration, not just publishing volume, is the lever that acts on recommendation instead of chasing a citation count that may already be downstream of a decision your content had no part in.
For B2B SaaS and other categories where the competitive field is genuinely crowded, the awareness-stage ghost rate matters most: that's where a model is most likely to cite you without recommending you, and where the gap between the two numbers is worth closing first. Run the co-occurrence check this quarter before the next reporting cycle locks in citation count as the headline number, and bring the split, not the blended total, into whatever meeting decides next quarter's content and PR budget.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.