AI visibility measurement is in the phase every new channel goes through, where the vendors publishing the data are also selling the tools, the methodologies are not comparable, and practitioners quote headline percentages at each other without checking what was counted. Paid search went through it. Social went through it. The way out is always the same: compare the datasets, find what survives the comparison, and build the standard from the overlap.
So we compared four. They differ in scale by four orders of magnitude, they measure different objects, and three of the four were published by companies with a commercial interest in the conclusion. That is not disqualifying. It is the normal condition of a young measurement discipline, and it is exactly why the agreement between them is more interesting than any individual figure.
Why AI visibility measurement is still broken
The core problem is that AI visibility has no equivalent of the impression. In search, an impression is a well-defined event: the page was rendered in a result set at a position. Everything downstream, click-through rate, average position, share of voice, is built on that primitive.
AI answers have no such primitive. A brand can appear as a linked source, as a named entity in prose, as both, or as an unattributed influence on an answer that mentions nobody. These are different events with different commercial value, and no industry convention yet says which of them counts as visibility. Every tool picks one, reports it as the number, and produces a figure that is not comparable to the tool next door.
Methodology: what we compared and how
We selected four publicly documented 2026 datasets that each measure a different layer of the same funnel, and normalized what could be normalized. No new data collection was performed for this review; every figure below is attributed to its original publisher and linked. Where a dataset's methodology limits what it can support, we say so rather than extending it.
| DATASET | PUBLISHER | SCALE | WHAT IT MEASURES | MAIN LIMITATION |
|---|---|---|---|---|
| AI Visibility Index | Semrush | 126M US prompts, 22 industries, Jan-Apr 2026 | Category-level brand visibility across four engines | Industry cuts are coarser than real competitive sets |
| Ghost citation study | Semrush and Kevin Indig | 3,981 appearances, 115 prompts, 14 countries, Jun 2026 | Whether an appearance is a link, a name, or both | Small prompt set; prompt selection drives results |
| LLM referral breakdown | Ahrefs | One site's full referral log | Which page types actually receive AI clicks | Single site with an atypical free-tool profile |
| Cold email benchmark | Instantly | Billions of sends, thousands of workspaces, 2025 | Outbound reply distribution by percentile | Self-reported platform data; different channel entirely |
The fourth is deliberately from an adjacent channel. It is included as a control: outbound email is a mature measurement discipline with the same structural property, a heavily skewed distribution reported as an average, and it shows what this category's numbers will look like once the definitions settle.
The control earns its place. Instantly's figures put the average cold email reply rate at 3.43%, the top quartile at 5.5%, and the top decile at 10.7%. A single average across that spread is nearly uninformative, and outbound practitioners learned years ago to read the percentile bands instead. AI visibility is currently at the stage outbound was at before that lesson landed: headline averages, quoted confidently, describing a distribution nobody has looked at. Expect the percentile framing to arrive here within a year, and consider adopting it early.
One methodological note on what this review is not. We did not attempt to reconcile the absolute numbers across datasets, because they are not measuring the same object and forcing them onto a common scale would manufacture precision that does not exist. A prompt in Semrush's index, an appearance in the Indig study, and a referral session in the Ahrefs log are three different units. What can legitimately be compared is the shape of each finding and whether the shapes agree. They do, which is the entire basis for the standard proposed at the end.
Finding 1: citation and mention are separate metrics
The Semrush and Indig study logged 3,981 domain appearances across ChatGPT, Gemini, Google AI Mode, and AI Overviews, and split them three ways. Ghost citations, where a source link exists but the brand is never named in the answer text, accounted for 61.7%. Mentions with no citation took 25.1%. Both together happened 13.2% of the time.
How brand appearances break down across four engines (Semrush and Kevin Indig, June 2026)
The engine-level split is sharper than the aggregate. ChatGPT cited 87% of the time and named the brand 20.7%. Gemini inverted it at 21.4% cited and 83.7% named. Semrush's larger index corroborates the structural difference from an independent direction: ChatGPT cites an average of 15 sources per response, Gemini an average of 3.
Two datasets, two methodologies, same conclusion. Citation and mention are not two views of one thing. They are produced by different parts of the generation process and they fail independently, which is why we treat ghost citations as a distinct measurement problem rather than a subcategory of citation tracking.
Finding 2: the cited page is not the destination page
Ahrefs traced its own LLM referral traffic to the pages that received it. Free tools took 36.45%, product pages 23.12%, and the homepage 20.41%, with more than 80% of all AI referral traffic concentrated in those three page types. Editorial content, the surface most AI visibility programs are built on, split the remainder with everything else.
| PAGE TYPE | SHARE OF AI REFERRAL TRAFFIC | TYPICAL ROLE IN A GEO PROGRAM |
|---|---|---|
| Free tools | 36.45% | Rarely built; largest single destination |
| Product pages | 23.12% | Rarely optimized for extraction |
| Homepage | 20.41% | Almost never treated as a citation surface |
| Everything else | ~20% | Where nearly all budget goes |
Set against finding 1, this produces the most actionable result in the comparison. Articles earn citations. Product and utility pages receive clicks. A program that measures article performance by AI referral sessions will conclude its best-performing asset is failing, because the article did its job and handed the visit to a different URL. We unpacked the operational consequences separately in AI search traffic by page type.
Two caveats belong here, both from Ahrefs' own write-up. The tracking is incomplete, since not all assistant traffic is recorded correctly, and AI Overviews traffic is lumped in with organic search. The site is also atypical: a large free-tool suite and a category-defining brand inflate the top two buckets relative to a normal B2B company. Treat the ordering as the finding and the percentages as theirs.
Finding 3: concentration sets your ceiling before tactics do
Semrush's index benchmarked 22 industries and found top-three concentration varying more than fortyfold in spread: 82.9% in news and media, 76.9% in consumer electronics, against 42.2% in industrial and 41.4% in finance. Only 36 brands across all 22 industries held visibility on every platform in every month of the study.
This is the variable most programs never measure and the one that best predicts whether a year of work will move anything. It also reframes the other findings: in a category at 83%, the marginal value of better extraction structure is small because the answer space is claimed. At 41%, the same work compounds against an unfinished field. The strategic reading is in AI search visibility by industry, and the practical instruction is simply to measure competitors' share alongside your own, which almost no tracking setup does by default.
Finding 4: almost nobody can measure any of this
The same index reports that 45% of marketing leaders cannot accurately measure brand visibility in AI answers, and only 9% have tools covering every relevant metric across platforms. It also found that organizations running fully integrated search and AI strategies reported increased traffic or leads 81% of the time, against 36% for those managing the two separately.
That 81-to-36 gap is the largest effect size anywhere in the four datasets, and it is organizational rather than tactical. It is worth more than any individual optimization discussed above, and it costs nothing but a reporting line and a reorganized team. Take that seriously before taking any tactic seriously.
“The largest measured effect in the entire comparison was not a tactic. It was whether search and AI were run by one team or two.”
Where the datasets disagree
Honest comparison means naming the contradictions, and there are three worth flagging before anyone builds a plan on the overlap.
The correct posture is to use these to calibrate expectations and to run the equivalent measurement on your own prompt set before committing budget. Every one of these findings is checkable on your own domain in under a week, and the version measured on your buyers' actual questions outranks any published benchmark.
There is a fourth disagreement that is really an absence. None of the four datasets measures downstream revenue. Semrush reports a conversion-value ratio elsewhere in its work, and the referral data records sessions, but no public dataset yet connects a specific citation to a specific closed deal. Every claim about the commercial value of AI visibility, including the ones we make, currently runs through proxy metrics. That gap is the honest limit of the entire discipline in 2026, and anyone presenting AI visibility as a directly attributable revenue channel is extrapolating past the available evidence.
We would rather state that plainly than paper over it. The case for investing here does not depend on precise attribution; it depends on the observation that buyers are demonstrably starting research in these surfaces and that being absent from an answer is not a neutral outcome. That argument is strong enough without borrowed precision, and programs built on it survive scrutiny better than programs built on a confident number that turns out to be a proxy.
The AI visibility measurement standard we recommend
The overlap between four incompatible methodologies supports a small, specific standard. This is what we now build into reporting and analytics engagements, and it is deliberately minimal.
Everything else is optional. Teams routinely build far more elaborate dashboards than this and still cannot answer whether they are winning, because they collapsed the two booleans into one score at the point of collection and cannot get them back. Collect both, blend never, and the rest of the analysis stays available. The upstream question of what earns a citation in the first place is a separate discipline, covered in the anatomy of an AI citation, and it only becomes answerable once the measurement above exists.
For teams starting from nothing, the sequencing matters more than the tooling. Build the prompt set and the two-column log first, run it manually for a month, and only then decide what to automate. We run this order in every audit engagement, and in technical categories it routinely reveals that the industry-level concentration figure badly overstated how contested the buyer-level questions actually were, which is what happened in a technical category build. For B2B SaaS specifically, the gap between category-level and buyer-level concentration is usually the whole opportunity.
Cite this research
Underlying sources, in order of scale: Semrush's 2026 AI Visibility Index for category concentration, engine citation counts, and the measurement-capability figures; the Semrush and Kevin Indig ghost citation study of June 2026 for the appearance breakdown; Ahrefs' LLM referral analysis for page-type distribution; and Instantly's 2026 cold email benchmark report for the control comparison on skewed distributions. No new data was collected for this review, and no figure above is ours unless explicitly labelled as such.
If you replicate any of this on your own prompt set, the number we would most like to see published is the one nobody has: an independent, non-vendor measurement of the ghost citation rate at scale. The 61.7% figure is doing an enormous amount of work in this industry right now, and it rests on 115 prompts. It deserves a bigger sample, and whoever runs it will have produced the most useful contribution to generative engine optimization measurement this year.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.