Monday morning, someone opens the AI visibility dashboard for the exec review, and the number is down twelve points from last week. Nobody touched the content. Nobody changed the target prompts. The engine just answered differently this time. Now somebody has to explain a swing that has nothing to do with the work anyone actually did — and that's precisely the scenario Wil Reynolds is warning teams about when he calls AI visibility metrics a vanity metric.
Why AI Visibility Metrics Feel Unstable
Every enterprise team tracking AI search visibility right now is running some version of the same setup: a list of prompts, checked against ChatGPT, Gemini, Perplexity, and whatever else, rolled into one trending line that says something like "AI mention rate: 34%, up 3 points." It gets screenshotted into a deck. It gets asked about in a QBR. And increasingly, it gets questioned by the person paying for the tool that produces it.
That questioning has a name attached to it now. Wil Reynolds, CEO of Seer Interactive, published a piece on his own company's blog on February 3, 2026 — "AI Visibility is a Vanity Metric, here's how to tell your boss" — that argues something sharper than the usual AI-hype-skepticism take. He's not just saying AI visibility is overhyped the way SEO rankings once were. He's saying AI search is worse for this than classic SEO ever was, because it lacks the one thing that made a ranking number trustworthy in the first place: a stable structure underneath it.
That's a different critique than the one we covered a few days ago, when Reynolds told Search Engine Journal that AI visibility "needs to be tied to an outcome." This piece isn't about outcomes at all. It's about whether the raw AI visibility metrics feeding your dashboard can be trusted to mean the same thing twice — and if you want the outcome-attribution argument, we already made our case on that one. This is a narrower fight, about measurement itself, and it deserves its own answer.
Wil Reynolds' Case Against AI Visibility Metrics
Reynolds' argument rests on three specific, checkable problems with how AI search behaves, and none of them are hand-waving. Take them one at a time, because each one lands differently.
“Reynolds' core claim, paraphrased from his Seer Interactive post: the same prompt run twice can return a different answer, a different number of sources, and a different set of citations — with no accounting yet for personalization. He cites Rand Fishkin's research showing AI responses are, in Fishkin's words, "NOT repeatable." A rank tracker can tell you rank #1 means a stable amount of traffic. Nothing in AI search holds still long enough to make that same promise.”
The second problem is structural. Classic SEO had a fixed shape: ten blue links, a known set of ranking factors, a relationship between position and click-through that held steady enough to build a forecast on. AI search has no fixed shape at all. Reynolds notes that the same prompt can return anywhere from three to seventeen results, and separately, that training cutoffs create a lag nobody's dashboard accounts for — his example is Gemini 3's February 2025 training cutoff, which means optimization work shipped in February 2026 simply won't show up in its answers yet, no matter how good it is. A number that can't distinguish "this isn't working" from "this hasn't been indexed into the model yet" isn't a number you can act on.
The third problem is about what visibility is even supposed to buy you. A client study Reynolds cites found that 44% of people include a specific brand name directly in their AI prompts. That means nearly half of the prompts feeding a typical visibility score aren't exploratory at all — they're brand-aware searches from people who already know who they're asking about. Showing up more often in that kind of prompt doesn't create discovery. It just confirms trust that was built somewhere else. His prescribed fix: stop chasing the AI-visibility-ranking dashboard, and report newsletter signups, content shares, direct traffic, and brand search volume instead — the things that actually demonstrate an audience believes you're worth paying attention to.
| REYNOLDS SAYS | THE CASE STUDIES SHOW | |
|---|---|---|
| What's unstable | Same prompt, different answer. Fishkin's research calls AI responses "NOT repeatable," result counts swing from 3 to 17, and training cutoffs hide months of real work. | Citation position and context on G2's own product pages held steady enough to attribute a 44% citation lift to one specific change. |
| What to measure instead | Newsletter signups, content shares, direct traffic, brand search volume. | Citation position on a narrow set of purchase-stage prompts, tracked per engine, not blended. |
| Where brand awareness cuts both ways | 44% of prompts already name the brand directly — visibility alone doesn't create discovery if trust wasn't already there. | 94% of B2B buyers consult an answer engine before talking to a salesperson, meaning even brand-aware prompts are live evaluation moments worth winning. |
What the Case Studies Show Instead
Here's the part of the argument Reynolds doesn't fully reckon with: some AI citation behavior is stable enough to build a program around, provided you stop measuring the wrong slice of it.
Profound's recap of its Zero Click New York 2026 event, published June 15, 2026, includes a detail that undercuts the "nothing about this holds still" framing without contradicting Reynolds' underlying data. G2 added context summaries to its product pages and saw a 44% increase in citations. That's not noise. That's a specific input, on a specific set of pages, followed by a specific and repeatable-enough output. If AI citation behavior were as chaotic as Reynolds' broader argument implies, a change like that wouldn't show up as a clean, attributable lift — it would show up as static.
The recap makes the more important claim explicit: citation position matters more than citation frequency for revenue signals. That's the distinction most AI visibility metrics collapse. A blended mention-rate score treats "we got named once, buried under six other sources" the same as "we're the source the engine leads with." Those are not the same event, and only one of them is a business signal.
The recap also flags LinkedIn as the most-cited domain for professional queries, with citations up 2x year-over-year — a domain-level signal, not a brand-level one, but it tells you where the engines are already pulling trusted answers from in B2B categories. And it states plainly that 94% of B2B buyers consult an answer engine before ever talking to a salesperson, with evaluation timelines compressing 40% as a result. Reynolds' 44%-brand-name-in-prompt figure and Profound's 94% figure aren't actually in conflict. They describe the same moment from two angles: buyers arrive at AI answers already partway informed, and the engine's answer is still shaping whether they shortlist you before a human ever gets involved.
The Verdict: Track Position on a Narrow List, Not Frequency on Every Prompt
Reynolds is right about the failure mode he's naming. A broad, blended AI visibility score — checked across every prompt you can think of, averaged into one weekly number — is exactly as unstable as he says. It will swing for reasons that have nothing to do with your content: a non-repeatable response, a result count that happened to land at four instead of eleven, a model that hasn't ingested your February work yet. Reporting that number to an executive as if it were a stable trend is a mistake, and teams that do it are going to get the number cut the first time someone in finance asks why it dropped for no reason anyone can name.
But throwing the whole category out and retreating to newsletter signups and brand search volume gives up something the G2 and LinkedIn data show is real: a narrower kind of AI visibility metric that doesn't share the instability problem, because it isn't trying to measure everything at once.
The distinction is this. Don't track how often you're mentioned across every plausible prompt a buyer might type — that's the noisy, aggregate number Reynolds is right to distrust. Instead, build a short, deliberately curated list of purchase-stage prompts, the kind someone types when they're two steps from a shortlist, not the kind someone types while browsing. Track citation position and citation context on that narrow list, per engine, the way G2's product-page work implicitly did. Frequency across everything is noise. Position on the prompts that matter is signal.
That's a more specific claim than "measure grounding, citation, and mention as three separate numbers," which is the case we made in our framework for measuring GEO grounding, citation, and mention as separate market-share numbers. That framework is about which event to measure. This one is about which prompts are even worth measuring in the first place, and what to look at once you've picked them — position and context, not raw frequency.
How to Build AI Visibility Metrics That Survive a Budget Review
We ran a version of this narrowing exercise with a patient engagement platform, where the mistake wasn't a lack of tracking — it was tracking too much. The team had a blended visibility score across dozens of loosely related prompts, and it moved around enough every week that leadership had stopped trusting it entirely. Cutting the tracked list down to a dozen purchase-stage prompts, and reporting citation position instead of mention frequency, turned a number nobody believed into one the CFO actually referenced in a renewal conversation.
Practically, that means three things in sequence. First, pull your current prompt list and mark which ones are actually close to a buying decision versus exploratory or top-of-funnel — most lists are 80% the latter, which is exactly why the aggregate score bounces around. Second, for the purchase-stage subset, track position and surrounding citation context per engine, not a blended average. Third, correlate that narrower number against pipeline over a real measurement window before you report it as causal — the discipline we lay out in our approach to attributing pipeline to organic and AI, and the same instinct behind our broader research into what actually earns an AI citation.
That's the sequencing work we run inside our reporting and analytics practice — not because narrowing the prompt list is the exciting part of the story, but because it's the part that determines whether the number on next quarter's slide survives contact with someone who's already skeptical of AI visibility metrics. Reynolds gave that skepticism a name. The G2 data gives it a boundary.
The dashboard that swings twelve points for no reason was never going to survive a hard question. The one built on a dozen prompts you chose on purpose, measuring position instead of noise, might be the first AI visibility number in your reporting stack that actually does.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.