I've sat in three QBRs this year where someone pulled up a dashboard, pointed at a number going up and to the right, and said some version of, 'See, our AI visibility is improving.' Every time, I asked which engine. Every time, there was a pause.
The number on the slide
That pause is the whole problem. The dashboard wasn't lying, exactly. It just wasn't built to answer the question anyone in the room actually cared about, which was never 'is our score up.' It was 'can a buyer asking ChatGPT about our category actually find us.' Those are not the same question, and 2026 has produced three separate studies that explain why.
None of this is a takedown of AI visibility tracking as a category. It's a warning about the version most teams bought first: one blended number, trending nicely, telling you almost nothing you can act on.
The one-number problem in AI visibility tracking
Rank tracking earned the right to be a single number because ranking is a single, stable event. A page holds position three for a keyword, or it doesn't. AI answers are not one event. A single buyer prompt can trigger five to twenty internal sub-retrievals inside an agentic system before the model settles on what to say, according to Michael King's May analysis at iPullRank, the same shift we walked through in what AI agents actually do when they get stuck. Compress that whole process into one visibility percentage, and you've thrown away the part of the data that would have told you what to fix.
A prototype is not a product. A blended score is not a measurement system. It just looks like one on a slide.
None of this is an argument for ignoring the number entirely, either. A rising blended score is still better than a falling one, all else equal. It's an argument for treating it the way you'd treat a single vanity metric anywhere else in marketing: fine as a headline for a slide, useless as the basis for a budget decision.
I'd go further: the teams most at risk here are the ones who did everything right on the reporting side by the old rules. They built a dashboard, they picked a metric, they trended it over time, exactly what a rigorous marketing org is supposed to do. The mistake wasn't sloppiness. It was applying a rank-tracking mental model to a measurement problem that doesn't behave like rank tracking, and nobody flagged the mismatch until research caught up with the tooling.
Ghost citations: a mention is not a citation
Growth Memo's April research, tracking 3,981 domains across 115 prompts, 14 countries, and 4 AI search engines, found a real gap between how often a brand gets mentioned in an AI answer and how often it gets formally cited with a link. Most GEO tools count both the same way. Some only track the easier-to-scrape mention, because it's cheaper to detect. A brand can look visible in that tool and be functionally invisible to a buyer who never sees a link to click, the exact gap our GEO readiness framework treats as a separate scored dimension rather than folding it into one number. Those posts about your rising visibility score aren't lying. They're leaving out the part where nobody could actually get to you.
The consensus gap: visibility doesn't travel
Here's the bigger one. Growth Memo's May study, what it calls the consensus gap, found that 91% of AI citations show up in exactly one search engine. Only 9% get cited by two or more engines for the same query.
Citation portability across AI search engines (Growth Memo, May 2026)
A blended mention-rate score cannot show you this. It can report a healthy average while you're, in reality, strong on one engine and a ghost everywhere else your buyers are actually asking. If your buying committee splits its research across ChatGPT, Perplexity, and Google AI Mode, the same way our own attribution work keeps finding they do, and 91% of citations don't carry between engines, a single score is measuring the wrong unit entirely. It's the search-visibility equivalent of tracking 'average temperature across all four seasons' and wondering why nobody can dress for the weather.
The undercount: what you see is a fraction
King's iPullRank analysis goes further and argues that the citations a tracking tool actually observes, the ones that make it into the final rendered answer, under-report a brand's real retrieval footprint by a factor of three to ten. Agentic systems research far more than they cite. Getting pulled into the research pass and never surfacing in the final answer still shapes what the model believes about your category, and no dashboard built to count visible citations sees any of it.
That's not a small asterisk. That's most of the iceberg.
And it's an iceberg most teams are sailing toward at speed. ConvertMate's 2026 GEO benchmark found 92% of marketers now say they're planning GEO investment, but only 40.6% have actually implemented anything. That gap between intent and execution is exactly where a misleading dashboard does the most damage: a team about to make its first real GEO investment case to leadership is the team most likely to lean on the one flattering number the tool hands them, instead of asking what it's hiding.
This isn't even the first time a supposedly authoritative dashboard turned out to be quietly wrong. Kevin Indig's own February research, analyzing 450 million impressions, found Google Search Console data itself is roughly 75% incomplete once you account for the queries and impressions Google filters out of the reporting UI. Marketers have spent a decade treating GSC as ground truth despite that gap. AI visibility tracking is repeating the same pattern faster: a young, unaudited measurement layer getting treated as gospel before anyone's checked what it's dropping.
What good AI visibility tracking looks like instead
Break mention rate out by engine. Not blended, not averaged, per engine, every time you report it. Separate true citations, the ones with a link, from bare mentions, the ones without. And treat any vendor's number as a floor, not a ceiling, given the retrieval you cannot see.
None of this requires new tooling most agencies don't already have access to. It requires refusing to report the one number the dashboard defaults to, and asking your GEO vendor, or the reporting team building your dashboard, for the breakdown instead. We've built our own citation research around exactly this split, engine by engine, mention versus citation, because the aggregate version was never going to survive contact with a real buying committee.
I'd also push back gently on the vendors themselves here, because most of them aren't hiding this on purpose. Building a per-engine, mention-versus-citation, retrieval-aware measurement product is genuinely harder and more expensive than shipping a blended score, and the market rewarded the simpler dashboard for years before this research existed to complicate the pitch. The tools aren't lying to you out of malice. They're lagging the research by the normal amount any tooling category lags the papers that eventually reshape it. That's still your problem to manage, today, with the budget conversation you have next quarter, but it's worth keeping the blame proportionate.
Ask your vendor a blunt question the next time they present the number: what does this score do when we're strong on one engine and invisible on three others? If the answer is 'it goes up,' you've found the flaw yourself, in the room, before it costs you a budget conversation.
One more thing worth adding to the breakdown: cross-reference citations against Ahrefs' finding that brand-mention volume correlates with AI visibility at 0.664, far ahead of raw backlink counts at 0.218. If your per-engine, mention-versus-citation numbers aren't moving, but your brand-mention volume across the open web is climbing, that's a leading indicator the lagging metric hasn't caught up to yet, not proof the strategy isn't working.
You're not behind if your mention-rate score looks flat this quarter. You might just be looking at a number that was never built to move the way you think it should. Ask for the breakdown before you decide the strategy failed.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.