I keep seeing this in QBRs. Someone pulls up a chart, points at a line trending up and to the right, and says some version of, 'Our AI visibility score is climbing.' I ask which engine. There's usually a pause, then, 'All of them. Blended.'
Rank tracking earned the right to be a single number because ranking was a single, stable fact. A page held position four for a keyword, or it didn't, and Google held roughly nine in ten searches for most of a decade, so the ground under that number barely moved either. Nobody had to ask which search engine, exactly, or whether that engine's importance had changed since Tuesday.
AI visibility tracking took the shape of that old dashboard without earning the same trust. It hasn't earned it yet, and it might not for a while. Two studies out this month explain why, and neither one is really about the number on your slide. They're both about what that number is quietly averaging away.
Start with the mechanism, because it's the one most dashboards don't even ask about. If two engines don't retrieve information the same way, averaging their citation behavior together tells you less than nothing. It tells you a number that feels like a fact and behaves like a rumor.
Claude doesn't work the way you assume ChatGPT works
Josh Blyskal at Profound published a July 22 breakdown of how Claude actually builds its answers, and the finding that should worry anyone running one dashboard across engines isn't really about citations. It's about search behavior itself.
When Claude did search the live web, 79.2% of what it cited traced back to Brave's top 10 results. That part sounds familiar. It's the kind of source-concentration story GEO teams already expect from ChatGPT and Perplexity. The number sitting right next to it is the one that changes the story: Claude only invoked a live web search at all for 36.6% of the prompts Profound tested.
Claude, by the numbers (Profound, July 22, 2026)
Read those two numbers together and the picture flips. Most of the time you're being 'tracked' against Claude, you're not being tracked against a search engine at all. You're being measured against whatever Claude already learned in training, or already has cached, and no amount of clean structure, machine access, or third-party corroboration reaches an answer that was never retrieved in the first place.
“Citation behavior and retrieval behavior are not the same fact. One describes what a search-triggered answer cites. The other describes how rarely the search gets triggered at all. A single Claude row on a dashboard has to pick one. Most tools pick whichever is easier to demo.”
It also means every hour spent on machine access, structure, or corroboration for Claude specifically is only reachable a little more than a third of the time. That's not a reason to skip the work. It's a reason to stop expecting it to move Claude's row on the dashboard as fast as it moves ChatGPT's.
This is really the same structural problem we documented in our review of how AI engines disagree on sources: the engines aren't running one process with different settings. They're running different processes. It also lines up with our own 2026 AI citation data review: pooling engines together erases exactly the signal a budget decision needs.
The market moved twice while your dashboard sat still
Kevin Indig's Growth Memo published its H1 2026 halftime report on July 27, and the topline number is worth sitting with: ChatGPT's share of AI-search activity dropped from 78% to 56% over six months. Gemini climbed to 30% over the same stretch. Claude rose to 10%. Google AI Mode, which barely registered as a line item a year ago, crossed 1 billion monthly active users.
Compare that to what SEO teams got used to for the better part of a decade. Google's share of desktop and mobile search barely moved in a given year, low single digits at most, which is exactly why an annual rank-tracking review was ever a defensible cadence in the first place.
That's not AI search market share drifting the way SEO market share used to drift, a point or two a year, safe to average over. That's a 22-point swing in the single most dominant engine, inside one half of one year.
A blended AI visibility score can absorb a swing like that and still climb. If your brand happens to be strong on the engine gaining share, or weak on the one losing it, the average can trend in exactly the wrong direction relative to what leadership assumes it means. Nobody in the room would know to ask.
Why one AI visibility tracking number is a category error
Put Blyskal's finding and Indig's finding next to each other and the case against a single score stops being about precision. Imprecise would be fine; every metric loses some resolution when you roll it up. This is worse than imprecise. It's a category error, the same mistake as averaging temperature across four seasons and reporting the result as 'the weather.'
A blended score assumes two things have to both be true to mean anything. First, that the systems underneath behave similarly enough that averaging preserves signal instead of destroying it. Second, that their relative weight holds still long enough for a trend line to mean what a trend line is supposed to mean. Blyskal's data says the first assumption is false at the mechanism level. Indig's data says the second is false at the market level, twice in six months.
Picture the actual budget meeting. A blended score is flat, or even slightly up, and the team reads that as steady progress worth funding at the same level. Underneath it, ChatGPT lost more than a fifth of its share and Gemini and Claude both grew. The budget that made sense in January is now funding the wrong mix of work, and the one chart in the room said everything was fine.
This builds directly on a complaint we've made before about the mention-rate dashboard most teams already run: that tool undercounts and misattributes even within a single engine. Layer engine-level volatility and mechanism-level differences on top of that, and a blended number isn't just missing detail. It's telling leadership the wrong thing happened, in either direction, depending on which way the mix moved underneath it.
The habit you inherited from rank tracking, and why it doesn't fit
None of this is a knock on the teams running one blended dashboard today. They built it the way a rigorous marketing org is supposed to build a metric: pick a number, trend it, defend the budget with it. That's exactly right for a market that behaves like the one rank tracking was built for. It's the wrong shape for this one.
Most generative engine optimization programs are still reporting against last year's engine assumptions, because last year's assumptions were the only ones available when the reporting got built. That's not a strategy failure. It's a timing problem, and 2026 is exposing it faster than most budget cycles can react to.
SEO budget cycles could run annually because the underlying market cooperated. GEO budget cycles inherited the same cadence without inheriting the same stability. Budget cycles built for SEO's slower pace don't survive AI engine volatility like this: a single half-year rearranged who matters most, more than once, before most teams had even scheduled their next formal review.
That's the actual mental-model swap, not a tooling upgrade. SEO budget cycles could treat the engine as a constant and argue about tactics. GEO budget cycles have to treat the engine mix itself as the variable, because per Growth Memo's numbers, it's moving faster than most teams' meeting cadence.
What AI visibility tracking should look like this quarter
The fix isn't five new dashboards nobody reads, and it isn't ignoring the blended number entirely; a rising average is still, all else equal, better than a falling one. The fix is refusing to let the blended number be the only thing leadership sees, and rebuilding the review cadence around a market that's already proven it won't hold still for a year. If your reporting and analytics function can't produce that breakdown today, that's the gap to close before the next budget cycle, not after.
None of the four moves above requires new tooling most agencies don't already have access to. It requires refusing to let the blended number be the default slide, and asking the team building your dashboard for the breakdown by name, every time.
None of this is an argument for giving up on measurement, or for standing up five dashboards nobody reads. It's an argument for admitting that 'AI visibility' was never one system, so it was never going to survive being reported as one number.
Your leadership team doesn't need a cleaner line going up and to the right. It needs to know that ChatGPT dropped 22 points of share in six months, and that Claude spends more of its time answering from memory than from the open web, because those two facts change where the next dollar goes. A single score can't tell them that. Four numbers, checked every quarter instead of every year, can.
That's not more work for the sake of more work. It's the same discipline SEO used to run, back when Google actually held still long enough to deserve a single number. It just doesn't hold still anymore. Neither does the market you're actually trying to measure.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.