Something Inc.Schedule a free consultation
GUIDE

How to Actually Measure GEO: Grounding, Citations, Mentions, and Market Share

Most GEO measurement reports one number and calls it visibility. This guide breaks it into the four things it actually is, and what to track for each.

INTERMEDIATE4 CHAPTERS17 MIN

Ask most marketing teams how their brand is doing in AI search and they'll show you one chart: a mention rate, or a citation count, trending up or down over the last month. It's a real number, and it's not wrong. It's also one measurement out of at least four genuinely different things a brand can track in generative search, and the other three matter just as much, sometimes more, for deciding whether a GEO program is actually working. This guide walks through each layer, in the order a fact actually moves through an AI engine: grounding, citation, and mention inside a single answer; then the separate blind spot most rank-tracking tools have around all three; then the fact that 'AI market share' itself is at least three different, non-interchangeable metrics being reported as if they were one. It ends with what an actual measurement stack should look like once you've separated all of it out.

Why one GEO measurement number isn't enough

The instinct to collapse GEO measurement into a single dashboard number is understandable. Leadership wants one chart, one trend line, one thing that's either going up or down. But a single blended number hides exactly the information that would tell you what to do next. A citation count that's flat could mean your content isn't good enough to be selected as a source, or it could mean an engine is grounding your pages constantly but rarely surfacing them as a numbered citation, which is a completely different problem with a completely different fix. A market-share report that shows one engine collapsing could mean your buyers are moving to a different platform, or it could mean the report is measuring app usage while your buyers are reaching you through referral clicks, a different number entirely. Every chapter in this guide exists because collapsing these distinctions into one metric has a real cost: it points teams at the wrong fix.

TL;DR · 60 SECONDSThis guide covers four distinct measurement layers: (1) grounding, citation, and mention are three separate events inside a single AI answer, not synonyms, per Dan Petrovic's framework; (2) traditional rank-tracking tools have a structural blind spot around AI Overview mentions specifically, which means a clean rank-tracker report can coexist with real, untracked AI-search visibility; (3) 'AI market share' is reported as at least three different metrics, audience/usage share, referral-traffic share, and AI-search query share, that move independently and shouldn't be compared directly; (4) a real measurement stack tracks all of the above separately, on a recurring cadence, rather than blending them into one chart.

Chapter 1: Grounding, citation, and mention are three different things

IN THIS CHAPTERA page can be grounded without being cited, cited without being mentioned by name, or mentioned without ever being formally cited. Each is a separate, independently trackable event.

Dan Petrovic, publishing on DEJAN's blog on July 7, 2026, laid out a distinction that most GEO reporting still doesn't reflect: grounding, citation, and mention are three separate events, not three names for the same thing. A grounding source is a page the model retrieves into its context window as raw material for generating an answer. On its own, Petrovic notes, it's invisible: the page contributed to the answer, but nothing on the surface tells you it did. A citation is a grounding source the model actively chose to surface as a reference, the numbered link a user can click through to. Petrovic cites a framing from Natzir Turrado that captures the distinction well: citations are the bibliography, not the brainstorm. And a mention is different again, the brand's name appearing directly in the generated answer text itself, the part a user actually reads, which Petrovic notes is driven more by the model's prior trained knowledge of the brand than by whatever was actually in the grounding set for that specific query.

EVENTWHAT IT MEANSIS IT VISIBLE TO THE USER?
GroundingPage retrieved into the model's context as raw materialNo, invisible by default
CitationA grounding source selected and surfaced as a numbered referenceYes, as a clickable source link
MentionBrand named directly in the generated answer textYes, in the text the user actually reads

The practical consequence: these three can move independently of each other, and a brand tracking only one of them is missing real signal. A page can be grounded constantly, pulled into context on nearly every relevant query, and still never get formally cited, because the engine chose other sources for the visible reference list. A brand can get cited, the numbered link present, without ever being mentioned by name in the prose, which matters because the anatomy of an AI citation shows citation rank and mention are read differently by users scanning an answer. And a brand can be mentioned by name with no citation at all attached, purely because the model's training data already associates the brand with the topic, independent of anything retrieved for that specific query.

This has a direct measurement implication. If your GEO tracking only counts citations, the numbered links, you're missing both the grounding layer (which tells you whether your content is even being retrieved, a prerequisite for everything downstream) and the mention layer (which tells you whether the model has actually learned to associate your brand with the topic independent of any single query's retrieval). A program that's grounded often but rarely cited has a content-selection problem: the material is there, it's just not winning the reference slot. A program that's cited but rarely mentioned by name has a different problem: the engine trusts the source enough to link it, but hasn't internalized the brand as a trained association yet. Those are different fixes, and a single blended 'citation rate' number can't tell you which one you have.

Here's a worked example of why the distinction matters in practice, not just in theory. Say a B2B software brand runs a weekly audit across its ten highest-intent buyer prompts. On six of them, its pricing and comparison pages get pulled into the model's context, grounding, every time. Of those six, only two result in a visible citation, a link the user could click. And across all ten prompts, the brand's name shows up directly in the answer text, a mention, on four of them, including two where no citation or confirmed grounding was logged at all. Read as one blended number, that's a confusing, mediocre-looking result. Read as three separate layers, it's an actionable one: the brand is being retrieved reliably (grounding is strong), but losing the reference slot to competitors most of the time (citation is weak), while still benefiting from some baseline name recognition the model picked up in training (mention exceeds citation). The fix that follows from that read, sharpen the comparison content specifically to win the citation slot, is not one you'd land on from a single mention-rate number alone.

Chapter 2: The blind spot your rank tracker has

IN THIS CHAPTERStandard rank-tracking tools were built to monitor traditional search results. AI Overviews and AI-mode answers can name a brand in ways those tools structurally can't see.

A widely discussed Hacker News thread in late July, pointing to a Medium piece on Google's AI Overviews recommending brands that traditional rank trackers never surface, put a name to a gap most GEO teams have run into without fully articulating it: your rank tracker was built to watch a specific thing, position for a specific keyword in a specific results layout, and an AI Overview or AI Mode answer doesn't necessarily present that way. A brand can be named favorably inside a synthesized answer, in a sentence a rank tracker has no mechanism to parse as a ranking event, while the tracker's dashboard shows nothing unusual at all.

This isn't a knock on rank-tracking tools broadly; they do what they were built to do, and position tracking for classic organic results is still a real, useful signal, covered in our own work on reporting and analytics. The gap is structural: most of these tools were designed around a results page with a stable, parseable layout, ten blue links or a small number of well-defined SERP features. An AI-generated answer is neither. It's a block of synthesized prose that can name a brand anywhere in its structure, in a sentence shape that varies query to query, engine to engine. Building a tool to reliably parse 'was this brand named, and how' out of that variability is a genuinely harder problem than parsing position ten of a stable ten-result page, which is part of why a lot of the tooling in this space is still catching up.

1A clean rank-tracker report doesn't mean clean AI-search visibilityDon't read 'no ranking issues flagged' as 'no AI-search issues exist.' The tools are watching different surfaces.
2Manual prompt auditing still catches what tools missRunning your top buyer queries directly through ChatGPT, AI Mode, and Perplexity by hand, on a recurring schedule, remains one of the most reliable ways to catch a mention or citation gap no automated tool surfaced yet.
3Purpose-built AI-visibility tools are maturing, but unevenlyA growing set of tools aim specifically at AI-answer tracking rather than classic rank tracking. Evaluate any of them against whether they can actually parse mention and citation events, not just whether they claim 'AI visibility' in their marketing.

The fix isn't to distrust rank tracking, it's to treat it as one input among several, not the single source of truth for AI-search visibility. A brand that's investing seriously in generative engine optimization needs a monitoring layer built specifically for AI answers, whether that's a purpose-built tool or a disciplined manual audit process, running alongside classic rank tracking rather than instead of it.

Chapter 3: Market share is at least three different numbers

IN THIS CHAPTER'AI market share' gets reported as one number, but at least three genuinely different things are being measured under that label, and they move independently.

This is the layer where the most public confusion happens, because three well-sourced, credible reports published within weeks of each other in mid-2026 all report 'AI market share' and land on meaningfully different pictures, because they're measuring different things.

46% / 28% / 10%
ChatGPT / Gemini / Claude 'true audience' usage share, Sensor Tower's State of AI 2026 (via Fast Company, Jun 16)
92.4%
ChatGPT's share of trackable LLM referral traffic, Previsible's 2026 State of AI Discovery Report
78% → 56%
ChatGPT's AI-search usage share, Jul 2025 to Jul 2026, Kevin Indig / Growth Memo

Sensor Tower's report measures 'true audience,' a cross-platform estimate of actual app and web usage, and finds ChatGPT already below 50%, at 46%, with Gemini at 28% and Claude at 10%. Previsible's report measures something narrower: referral traffic specifically, the share of trackable sessions arriving at websites from LLM referrals, and finds ChatGPT dominating that specific channel at 92.4%. Kevin Indig's Growth Memo tracks a third thing again: AI-search usage share, the share of AI-search-specific queries an engine handles, finding ChatGPT's share there falling from 78% to 56% over the same period Sensor Tower's audience number and Previsible's referral number were measured.

METRICWHAT IT ACTUALLY MEASURESSOURCE
"True audience" usage shareCross-platform app/web usage reachSensor Tower, State of AI 2026
Referral traffic shareShare of trackable website sessions referred from LLMsPrevisible, State of AI Discovery Report
AI-search usage shareShare of AI-search-specific query volume by engineKevin Indig / Growth Memo
None of these three numbers is wrong. They're measuring three different things, and the industry keeps reporting all three under the same label: market share.

None of these reports contradicts the others once you separate what each is actually counting. ChatGPT can simultaneously be losing ground in overall audience reach (Sensor Tower), still dominate the specific channel of referral traffic to websites (Previsible), and be losing share of AI-search query volume specifically (Indig), because 'has the most total users,' 'sends the most website traffic,' and 'handles the largest share of AI-search-specific queries' are three different competitions. A brand deciding where to invest GEO effort needs to know which of the three actually matters for its own buyers, not just which number showed up in the most recent report someone forwarded around.

A practical way to decide which market-share metric actually matters for your program is to work backward from the decision you're trying to make. If the question is 'where should our next dollar of GEO content investment go,' AI-search usage share, Indig's metric, is the closer proxy, since it tracks where AI-search-specific demand is concentrating. If the question is 'is our website traffic strategy exposed to a single engine,' Previsible's referral-traffic share is the more relevant number, since it's measuring the thing that actually shows up as sessions in your own analytics. If the question is 'how big is the overall addressable audience across AI products, full stop,' Sensor Tower's audience metric is the right one, but it's also the least directly actionable of the three for a content or GEO team, since audience size doesn't tell you anything about whether that audience is asking the kinds of questions your content answers.

The mistake to avoid is treating any one of these three as a scorecard for whether GEO is 'working.' A brand can execute a strong GEO program and see Previsible's referral-share number for its primary engine fall, not because the program failed, but because engine-level referral concentration shifted for reasons entirely outside any individual brand's content or authority decisions. Isolate the metric that matches the decision at hand, and report movement in the other two as market context, not as a verdict on program performance.

Chapter 4: Building the GEO measurement stack that actually holds up

IN THIS CHAPTERPut the three prior chapters together into a single, practical tracking structure, one that separates each layer instead of blending them.

A measurement stack built from the previous three chapters looks less like one dashboard and more like four coordinated views, each answering a distinct question, reported on its own cadence rather than merged into a single trend line.

Weekly
Grounding and citation trackingTrack whether your pages are being retrieved (grounding) and whether they're winning the visible reference slot (citation), separately, per engine.
Weekly
Mention tracking, independent of citationRun a recurring manual or tool-assisted audit of whether your brand is named in AI answers at all, regardless of whether a formal citation is attached.
Weekly
AI-answer visibility, alongside classic rank trackingKeep rank tracking for what it's good at, and layer a purpose-built AI-answer audit on top, precisely because the two tools have different blind spots.
Monthly
Market-share context, by the metric that matches your goalTrack the specific market-share metric relevant to your program (usage share, referral share, or query share), and don't blend it with the other two when reporting to leadership.

This is also where category-level durability fits in, a topic we've covered in depth elsewhere: once a brand becomes the established source for a specific topic, that ownership tends to hold far more reliably than engine-level market share does. That's a fifth, longer-cadence layer worth adding once the four weekly and monthly layers above are running: a quarterly check on which of your priority topics have a clear owner yet, and whether it's you.

It's worth being explicit about staffing and cadence, since 'track four separate layers' can sound like more overhead than it actually is once it's built into an existing routine rather than treated as a new project. The weekly grounding, citation, and mention audits are a single, roughly hour-long block once a query list and logging template exist, not four separate research efforts. The monthly market-share review is a matter of pulling the specific metric relevant to your program from whichever source publishes it, Growth Memo, Previsible, or Sensor Tower, and logging the number alongside your own referral and citation data rather than treating it as a standalone headline. Most teams already have someone doing a version of the weekly audit informally; the change this guide argues for is making it structured, logged, and separated by layer, instead of an ad hoc check that ends up blended into memory rather than a comparable dataset.

None of these four layers requires exotic tooling to start. Manual prompt audits across ChatGPT, Perplexity, AI Mode, and Claude, run on a fixed weekly schedule and logged consistently, cover the grounding, citation, and mention layers well enough to start acting on. What requires more discipline is keeping the layers separate once you start reporting them, resisting the pull toward one blended chart, because the moment they're merged back into a single number, you've recreated the exact problem this guide opened with.

Pair this measurement structure with the practical funnel view we've published separately: a citation isn't the same as a session, and crawl access, citation, and traffic each leak independently. The measurement layers in this guide sit upstream of that funnel, describing what happens inside a single answer and inside the reported market data; the funnel research describes what happens after, whether a citation actually turns into a visit. Run both frameworks together, and a GEO report stops being one number that leadership either likes or doesn't, and starts being an actual diagnostic tool: something that tells you specifically what to fix, because you can finally see which of five or six genuinely different things moved.

One last piece of discipline worth building in from day one: version and date every log entry. Engines change how they present answers, what counts as a citation, and how aggressively they ground versus mention, often without any public announcement. A measurement stack that doesn't timestamp its own methodology changes, a new query list, a new logging template, a newly available tool, will eventually produce a trend line that looks like a real shift in visibility but is actually an artifact of how the measurement itself changed. The pattern of unconfirmed ranking volatility we've tracked elsewhere this year is a useful reminder that even well-resourced third-party trackers routinely disagree with each other and with the platforms themselves; a smaller, internal measurement program should hold itself to the same honesty about what it can and can't confirm, rather than reporting every month-over-month wiggle as a confirmed trend.

Start small if a full four-layer buildout feels like too much at once. Pick your ten highest-intent buyer prompts, run them by hand across your priority engines this week, and log grounding, citation, and mention separately for each one, even in a plain spreadsheet. That single exercise, repeated weekly, will surface more real signal in a month than another quarter of watching one blended citation-count chart tick up and down without knowing why.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

TT
Tyler TruffiMANAGING PARTNER, SOMETHING INC.

Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.