Something Inc.Schedule a free consultation
ANALYTICS

Ghost citations: cited by AI, never named

In a Semrush and Kevin Indig study, 61.7% of brand appearances were links with no brand name in the answer text. Your visibility number is measuring the wrong half.

JBJosh BernsteinManaging Partner · AUG 13, 2026 · 11 MIN READ
61.7%
of brand appearances are ghost citations
25.1%
are mentions with no citation at all
13.2%
are both cited and named
TL;DR · 60 SECONDSSemrush and Kevin Indig looked at 3,981 domain appearances across ChatGPT, Gemini, Google AI Mode, and AI Overviews. Only 13.2% of the time did a brand get both a link and a name in the answer. Ghost citations, where the link exists but the reader never sees who wrote it, made up 61.7%. If your AI visibility report tracks one of those numbers, it is describing a minority of your actual presence.

There is a failure mode in AI visibility reporting that almost every team has and almost nobody has named. You are in the answer. The engine used your page, linked it as a source, and built its response partly out of your work. And the reader finished reading without ever learning your company exists.

That is a ghost citation, and new data suggests it is the normal case rather than the exception. In a June 2026 study, Semrush and Kevin Indig ran 115 prompts across 14 countries and logged 3,981 domain appearances across four engines. Ghost citations, defined as a source link appearing while the brand name never appears in the answer text, accounted for 61.7% of them. Mentions with no citation took another 25.1%. Both together, the outcome every program is implicitly optimizing for, happened 13.2% of the time.

What ghost citations are

A citation is a machine-readable fact: the engine attached your URL to its answer. A mention is a human-readable fact: the reader saw your name. They come from different parts of the generation process, and they fail independently. Ghost citations are what you get when retrieval works and attribution does not.

OUTCOMESHARE OF APPEARANCESWHAT THE READER EXPERIENCES
Ghost citation (cited, not named)61.7%A footnote link they probably do not click
Mention only (named, not cited)25.1%Your name in prose with no way to reach you
Both cited and named13.2%The outcome everyone assumes they are buying

Read the middle row again, because it is the stranger of the two failures. A quarter of the time the engine names a brand from its own parametric memory without linking anything. There is no click available. Nothing appears in your analytics. That appearance is invisible to every tracking method that starts from a referral, which is most of them, and it is one reason we keep arguing that the AI visibility reporting gap is a measurement problem before it is a content problem.

THE DISTINCTION THAT MATTERSCitations are earned by being retrievable. Mentions are earned by being nameable: an unambiguous entity the model can attach a claim to. Optimizing extraction improves the first and does almost nothing for the second.

The engine split nobody accounts for

The aggregate hides the most actionable finding in the study. ChatGPT and Gemini behave close to opposite, and averaging them produces a number that describes neither.

ChatGPT: cited87%
ChatGPT: named21%
Gemini: cited21%
Gemini: named84%

How often a brand appearance includes a citation vs a name, by engine

ChatGPT cited 87% of the time and named the brand 20.7% of the time. Gemini inverted it: 21.4% cited, 83.7% named. Semrush's larger index of 126 million prompts points at the same structural difference from another angle, reporting ChatGPT citing an average of 15 sources per response against Gemini's 3. One engine is a bibliography. The other is a conversation that happens to remember brands.

That has a direct planning consequence. If your buyers research in ChatGPT, your ceiling is citations and your job is to be retrievable and correctly attributed in the source list. If they research in Gemini, citations are scarce and your job is to be the entity the model recalls unprompted, which is a brand and corroboration problem rather than a page-structure one. Running one strategy across both and reporting a blended visibility score is how programs end up unable to explain their own numbers.

A blended AI visibility score across engines that disagree this violently is not a metric. It is an average of a bibliography and a conversation.

Why ghost citations break your reporting

Three specific distortions follow from tracking one number when the underlying reality has two axes.

1You undercount ChatGPT and overcount GeminiA citation-based tracker makes ChatGPT look like your best channel and Gemini look broken. A mention-based tracker reverses it. Both are artifacts of the counting method, and teams reallocate real budget on the strength of them.
2Content wins look like content failuresA page that gets cited constantly but never named will show strong citation counts and no brand lift, so it reads as an underperformer in any dashboard that weights recall. It is doing exactly what a source is supposed to do.
3You cannot attribute the pipeline you do winWhen 25.1% of appearances are names with no link, the buyer who arrives later types your company into a browser directly. That lands in direct or branded search. The AI channel that created the demand gets none of the credit, and the budget follows the credit.

The third one is the expensive one. It is the mechanism by which a working AI program gets defunded: demand created in an untracked surface, converted through a tracked one, and attributed to the tracked one. Anyone who ran brand-versus-performance attribution arguments a decade ago has seen this movie.

Five plays that turn ghost citations into named mentions

Ghost citations are not a defect to eliminate. A citation is a real asset. The goal is to raise the share that also carry a name, and the levers are more specific than publish better content.

1Your pages are cited but the answer credits nobodyPut the attribution inside the quotable sentence
THE MOVES
Write the key claim so the brand is grammatically inseparable from it
Prefer our 2026 analysis found X over research shows X
Name the company in the sentence that carries the statistic, not the paragraph before it
Repeat the attribution once mid-page so a mid-document chunk still carries it
DONE WHENThe sentence an engine would lift cannot be quoted without naming you.
2The engine names competitors instead of youPublish comparative content, not informational content
THE MOVES
Build head-to-head and alternatives pages for your category
State the comparison verdict plainly rather than hedging across options
Include a criteria table the model can lift as a unit
DONE WHENComparative pages exist for the ten prompts your buyers actually run.
3You are invisible in short conversational promptsOptimize for the short query, not the long one
THE MOVES
Collect the two-to-five word prompts buyers use, not the paragraph-long ones
Make sure a single page answers each one without qualification
Check whether your brand is nameable from that page alone
DONE WHENShort-prompt coverage is tracked separately from long-prompt coverage.
4Gemini never surfaces youFeed parametric memory, not just retrieval
THE MOVES
Pursue third-party corroboration on the sources models trained on
Keep entity data consistent across Wikipedia, Wikidata, G2, and Crunchbase
Budget earned media against engine visibility, not against PR reach
DONE WHENYour entity description is identical across every third-party source you control.
5Nobody can tell whether any of this workedSplit the metric before you optimize it
THE MOVES
Log citation rate and mention rate as two columns per engine
Freeze a prompt set so month-over-month comparisons are valid
Report the both-cited-and-named share as the headline number
DONE WHENEvery visibility report states which of the two numbers moved.

Play two has the strongest evidence behind it. The same study found comparative content produced 2.4x more brand mentions than informational content, which is consistent with what we have seen for a while: comparison pages earn the most citations and, it turns out, the most names as well. The buyer's first question is a comparison, and the answer to a comparison has to identify who is being compared.

The short-query finding in play three is the one most likely to change your roadmap. Short conversational prompts produced 30x to 50x more brand mentions than long ones. Long, heavily-specified prompts push the model toward synthesis, where names dissolve into a general answer. Short prompts push it toward recall, where names are the answer. Most enterprise prompt-tracking sets are full of long prompts because they read as more realistic, and they are systematically measuring the condition where brands are least likely to be named.

How to measure both numbers properly

The instrumentation is not complicated. It is just two columns where most teams keep one.

Minimum viable ghost citation tracking● LIVE
For each prompt in a frozen set, per engine, record:
 
cited = brand URL appears in the source list (0/1)
mentioned = brand name appears in the answer text (0/1)
 
Then report four rates, never one blended score:
 
ghost_rate = cited AND NOT mentioned
mention_only = mentioned AND NOT cited
full_rate = cited AND mentioned <- the headline
absent_rate = neither
 
Segment every rate by engine. Never average ChatGPT with Gemini.
Keep the prompt set frozen; changing prompts resets the baseline.

Report full_rate as the headline and the other three as diagnostics. A rising ghost rate with a flat full rate means your retrieval work is landing and your attribution work is not, which is a copy fix. A rising mention-only rate means parametric memory is improving, which is usually earned media finally compounding. This is the shape of the dashboards we build during reporting and analytics engagements, and the two-column split is the part clients keep after everything else changes.

One caveat on the numbers themselves: 115 prompts across 3,981 appearances is a real study but a modest one, and prompt selection drives results heavily in this kind of work. Treat 61.7% as directionally sound rather than precise, and run the same split on your own prompt set before you rebuild a roadmap around it. Semrush's own reporting notes that 45% of marketing leaders cannot accurately measure brand visibility in AI answers and only 9% have tools covering every relevant metric, so the honest position for most teams is that they do not yet know their own ghost rate.

Start here

Take twenty prompts your buyers actually run, weighted toward short ones. Run them across ChatGPT and Gemini. For each, mark two boxes: cited, and named. It takes an afternoon and produces the only version of this data that matters, which is yours.

Then look at the ghost rate specifically. If it is high, you have already won the hard half. Being retrieved is the part that depends on crawlability, structure, and trust, and the fix for the remaining half is largely editorial: write the claims so your name travels with them. If your ghost rate is low and your absent rate is high, the work is upstream, and it starts with the signals that earn a citation at all rather than with attribution phrasing. The full methodology behind the study is worth reading in Semrush's write-up of the AI visibility index, particularly if you operate in a category where the top three brands already hold most of the visibility. For B2B SaaS teams that concentration is the real competitive picture, and ghost citations are how you discover you are closer to it than your dashboard suggests.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

JB
Josh BernsteinMANAGING PARTNER, SOMETHING INC.

Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.