Something Inc.Schedule a free consultation
RESEARCH

AI visibility measurement: what four 2026 datasets agree on

We compared four public datasets covering 126 million prompts, 3,981 brand appearances, one site's full referral log, and billions of outbound sends. They disagree on almost everything except the thing that matters most.

RESEARCH4 DATASETS2026
126M
AI prompts in the largest dataset compared
61.7%
of brand appearances are citations without a name
9%
of organizations have tooling for every relevant metric
TL;DR · 60 SECONDSFour independent 2026 datasets measure AI visibility in four incompatible ways, and their headline numbers cannot be reconciled. But every one of them, approached from a different direction, lands on the same structural conclusion: AI visibility is at least two metrics, not one, and any program reporting a single blended score is describing something that does not exist. This is a comparative review of the four, what each can and cannot support, and the measurement standard the overlap implies.

AI visibility measurement is in the phase every new channel goes through, where the vendors publishing the data are also selling the tools, the methodologies are not comparable, and practitioners quote headline percentages at each other without checking what was counted. Paid search went through it. Social went through it. The way out is always the same: compare the datasets, find what survives the comparison, and build the standard from the overlap.

So we compared four. They differ in scale by four orders of magnitude, they measure different objects, and three of the four were published by companies with a commercial interest in the conclusion. That is not disqualifying. It is the normal condition of a young measurement discipline, and it is exactly why the agreement between them is more interesting than any individual figure.

Why AI visibility measurement is still broken

The core problem is that AI visibility has no equivalent of the impression. In search, an impression is a well-defined event: the page was rendered in a result set at a position. Everything downstream, click-through rate, average position, share of voice, is built on that primitive.

AI answers have no such primitive. A brand can appear as a linked source, as a named entity in prose, as both, or as an unattributed influence on an answer that mentions nobody. These are different events with different commercial value, and no industry convention yet says which of them counts as visibility. Every tool picks one, reports it as the number, and produces a figure that is not comparable to the tool next door.

THE MEASUREMENT PROBLEM IN ONE LINEIn search, everyone agreed what an impression was before they argued about how to get more of them. In AI search, the argument about tactics started first, and the definition is still unsettled.

Methodology: what we compared and how

We selected four publicly documented 2026 datasets that each measure a different layer of the same funnel, and normalized what could be normalized. No new data collection was performed for this review; every figure below is attributed to its original publisher and linked. Where a dataset's methodology limits what it can support, we say so rather than extending it.

DATASETPUBLISHERSCALEWHAT IT MEASURESMAIN LIMITATION
AI Visibility IndexSemrush126M US prompts, 22 industries, Jan-Apr 2026Category-level brand visibility across four enginesIndustry cuts are coarser than real competitive sets
Ghost citation studySemrush and Kevin Indig3,981 appearances, 115 prompts, 14 countries, Jun 2026Whether an appearance is a link, a name, or bothSmall prompt set; prompt selection drives results
LLM referral breakdownAhrefsOne site's full referral logWhich page types actually receive AI clicksSingle site with an atypical free-tool profile
Cold email benchmarkInstantlyBillions of sends, thousands of workspaces, 2025Outbound reply distribution by percentileSelf-reported platform data; different channel entirely

The fourth is deliberately from an adjacent channel. It is included as a control: outbound email is a mature measurement discipline with the same structural property, a heavily skewed distribution reported as an average, and it shows what this category's numbers will look like once the definitions settle.

The control earns its place. Instantly's figures put the average cold email reply rate at 3.43%, the top quartile at 5.5%, and the top decile at 10.7%. A single average across that spread is nearly uninformative, and outbound practitioners learned years ago to read the percentile bands instead. AI visibility is currently at the stage outbound was at before that lesson landed: headline averages, quoted confidently, describing a distribution nobody has looked at. Expect the percentile framing to arrive here within a year, and consider adopting it early.

One methodological note on what this review is not. We did not attempt to reconcile the absolute numbers across datasets, because they are not measuring the same object and forcing them onto a common scale would manufacture precision that does not exist. A prompt in Semrush's index, an appearance in the Indig study, and a referral session in the Ahrefs log are three different units. What can legitimately be compared is the shape of each finding and whether the shapes agree. They do, which is the entire basis for the standard proposed at the end.

Finding 1: citation and mention are separate metrics

The Semrush and Indig study logged 3,981 domain appearances across ChatGPT, Gemini, Google AI Mode, and AI Overviews, and split them three ways. Ghost citations, where a source link exists but the brand is never named in the answer text, accounted for 61.7%. Mentions with no citation took 25.1%. Both together happened 13.2% of the time.

Cited but not named62%
Named but not cited25%
Both cited and named13%

How brand appearances break down across four engines (Semrush and Kevin Indig, June 2026)

The engine-level split is sharper than the aggregate. ChatGPT cited 87% of the time and named the brand 20.7%. Gemini inverted it at 21.4% cited and 83.7% named. Semrush's larger index corroborates the structural difference from an independent direction: ChatGPT cites an average of 15 sources per response, Gemini an average of 3.

Two datasets, two methodologies, same conclusion. Citation and mention are not two views of one thing. They are produced by different parts of the generation process and they fail independently, which is why we treat ghost citations as a distinct measurement problem rather than a subcategory of citation tracking.

WHAT THIS RULES OUTAny single-number AI visibility score computed across engines. Averaging an engine that cites 15 sources with one that cites 3 produces a figure that describes neither, and it moves when the engine mix shifts rather than when your visibility does.

Finding 2: the cited page is not the destination page

Ahrefs traced its own LLM referral traffic to the pages that received it. Free tools took 36.45%, product pages 23.12%, and the homepage 20.41%, with more than 80% of all AI referral traffic concentrated in those three page types. Editorial content, the surface most AI visibility programs are built on, split the remainder with everything else.

PAGE TYPESHARE OF AI REFERRAL TRAFFICTYPICAL ROLE IN A GEO PROGRAM
Free tools36.45%Rarely built; largest single destination
Product pages23.12%Rarely optimized for extraction
Homepage20.41%Almost never treated as a citation surface
Everything else~20%Where nearly all budget goes

Set against finding 1, this produces the most actionable result in the comparison. Articles earn citations. Product and utility pages receive clicks. A program that measures article performance by AI referral sessions will conclude its best-performing asset is failing, because the article did its job and handed the visit to a different URL. We unpacked the operational consequences separately in AI search traffic by page type.

Two caveats belong here, both from Ahrefs' own write-up. The tracking is incomplete, since not all assistant traffic is recorded correctly, and AI Overviews traffic is lumped in with organic search. The site is also atypical: a large free-tool suite and a category-defining brand inflate the top two buckets relative to a normal B2B company. Treat the ordering as the finding and the percentages as theirs.

Finding 3: concentration sets your ceiling before tactics do

Semrush's index benchmarked 22 industries and found top-three concentration varying more than fortyfold in spread: 82.9% in news and media, 76.9% in consumer electronics, against 42.2% in industrial and 41.4% in finance. Only 36 brands across all 22 industries held visibility on every platform in every month of the study.

82.9%
top-three share in the most concentrated category
41.4%
top-three share in the least concentrated
36
brands with unbroken cross-platform visibility

This is the variable most programs never measure and the one that best predicts whether a year of work will move anything. It also reframes the other findings: in a category at 83%, the marginal value of better extraction structure is small because the answer space is claimed. At 41%, the same work compounds against an unfinished field. The strategic reading is in AI search visibility by industry, and the practical instruction is simply to measure competitors' share alongside your own, which almost no tracking setup does by default.

Finding 4: almost nobody can measure any of this

The same index reports that 45% of marketing leaders cannot accurately measure brand visibility in AI answers, and only 9% have tools covering every relevant metric across platforms. It also found that organizations running fully integrated search and AI strategies reported increased traffic or leads 81% of the time, against 36% for those managing the two separately.

That 81-to-36 gap is the largest effect size anywhere in the four datasets, and it is organizational rather than tactical. It is worth more than any individual optimization discussed above, and it costs nothing but a reporting line and a reorganized team. Take that seriously before taking any tactic seriously.

The largest measured effect in the entire comparison was not a tactic. It was whether search and AI were run by one team or two.

Where the datasets disagree

Honest comparison means naming the contradictions, and there are three worth flagging before anyone builds a plan on the overlap.

1Engine importance is unresolvedThe referral data implies ChatGPT dominates traffic. The appearance data implies Gemini dominates naming. Both can be true, and which matters more depends entirely on whether your funnel needs a click or a recollection. No dataset here settles it, and any tool claiming to has picked a side quietly.
2Sample sizes differ by four orders of magnitude126 million prompts and 115 prompts are not the same kind of evidence. The 61.7% ghost citation figure comes from the small study, and it is the number most widely quoted. It is directionally credible and precision-limited, and it should be described that way.
3Three of four publishers sell the remedySemrush, Ahrefs, and Instantly all sell tooling in the categories their data makes urgent. That does not make the numbers wrong, and the methodologies are disclosed. It does mean an independent replication would be worth more than another vendor study, and none exists yet.

The correct posture is to use these to calibrate expectations and to run the equivalent measurement on your own prompt set before committing budget. Every one of these findings is checkable on your own domain in under a week, and the version measured on your buyers' actual questions outranks any published benchmark.

There is a fourth disagreement that is really an absence. None of the four datasets measures downstream revenue. Semrush reports a conversion-value ratio elsewhere in its work, and the referral data records sessions, but no public dataset yet connects a specific citation to a specific closed deal. Every claim about the commercial value of AI visibility, including the ones we make, currently runs through proxy metrics. That gap is the honest limit of the entire discipline in 2026, and anyone presenting AI visibility as a directly attributable revenue channel is extrapolating past the available evidence.

We would rather state that plainly than paper over it. The case for investing here does not depend on precise attribution; it depends on the observation that buyers are demonstrably starting research in these surfaces and that being absent from an answer is not a neutral outcome. That argument is strong enough without borrowed precision, and programs built on it survive scrutiny better than programs built on a confident number that turns out to be a proxy.

The AI visibility measurement standard we recommend

The overlap between four incompatible methodologies supports a small, specific standard. This is what we now build into reporting and analytics engagements, and it is deliberately minimal.

Foundation
Freeze the prompt setThirty to fifty prompts drawn from real buyer questions, weighted toward short conversational phrasings. Never change them mid-quarter. A changed prompt set resets the baseline and every trend line built on it becomes meaningless.
Core metric
Record two booleans, never one scoreFor each prompt and engine, log cited and mentioned separately. Report the both-cited-and-named rate as the headline and the other three combinations as diagnostics. This single change resolves most of the contradictions above.
Core metric
Segment by engine, alwaysNever average ChatGPT with Gemini. Their citation behaviour differs by a factor of five, so a blended number moves with engine mix rather than with performance.
Strategic
Log competitors, not just yourselfRecord every brand appearing alongside you. Top-three share of total appearances is your category concentration, and it determines whether a broad program or a narrow one is the right call.
Attribution
Split citation surface from destination surfaceTrack which pages get cited and which pages receive AI referral clicks as two separate reports. Judging article performance by referral sessions is the single most common way a working program gets defunded.
The minimum viable schema● LIVE
prompt_id # stable across quarters; never renumber
engine # chatgpt | gemini | ai_mode | ai_overviews
run_date
cited # 0/1 brand URL present in source list
mentioned # 0/1 brand name present in answer text
cited_url # which page the engine linked
competitors[] # every other brand named or cited
 
# Derived, reported per engine and never blended:
# full_rate = cited AND mentioned <- headline
# ghost_rate = cited AND NOT mentioned
# mention_only = mentioned AND NOT cited
# concentration = top-3 competitor share of all appearances

Everything else is optional. Teams routinely build far more elaborate dashboards than this and still cannot answer whether they are winning, because they collapsed the two booleans into one score at the point of collection and cannot get them back. Collect both, blend never, and the rest of the analysis stays available. The upstream question of what earns a citation in the first place is a separate discipline, covered in the anatomy of an AI citation, and it only becomes answerable once the measurement above exists.

For teams starting from nothing, the sequencing matters more than the tooling. Build the prompt set and the two-column log first, run it manually for a month, and only then decide what to automate. We run this order in every audit engagement, and in technical categories it routinely reveals that the industry-level concentration figure badly overstated how contested the buyer-level questions actually were, which is what happened in a technical category build. For B2B SaaS specifically, the gap between category-level and buyer-level concentration is usually the whole opportunity.

Cite this research

CITATIONSomething Inc. (2026). AI visibility measurement: what four 2026 datasets agree on. A comparative review of four public datasets covering 126 million AI prompts, 3,981 brand appearances, one site's complete LLM referral log, and billions of outbound email sends. Available at somethinginc.com/blog/ai-visibility-measurement-four-datasets.

Underlying sources, in order of scale: Semrush's 2026 AI Visibility Index for category concentration, engine citation counts, and the measurement-capability figures; the Semrush and Kevin Indig ghost citation study of June 2026 for the appearance breakdown; Ahrefs' LLM referral analysis for page-type distribution; and Instantly's 2026 cold email benchmark report for the control comparison on skewed distributions. No new data was collected for this review, and no figure above is ours unless explicitly labelled as such.

If you replicate any of this on your own prompt set, the number we would most like to see published is the one nobody has: an independent, non-vendor measurement of the ghost citation rate at scale. The 61.7% figure is doing an enormous amount of work in this industry right now, and it rests on 115 prompts. It deserves a bigger sample, and whoever runs it will have produced the most useful contribution to generative engine optimization measurement this year.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

TT
Tyler TruffiMANAGING PARTNER, SOMETHING INC.

Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.