Picture the pitch. A vendor walks your team through a dashboard, and there's one number at the top of it — a friendly, single score that's supposed to tell you how visible your brand is across ChatGPT, Perplexity, Google AI Mode, all of it, boiled down into one line. The line goes up, the room nods. The line goes down, someone gets a Slack message before lunch. Almost nobody in that room can tell you what actually moved it. That's the quiet problem sitting underneath a big chunk of the AI visibility platform category right now, and a genuinely credible person just said so out loud, in public, by asking the entire industry to help him find out.
The Survey Question That's the Real Story
On July 14, 2026, Duane Forrester published a post on his newsletter, Duane Forrester Decodes, with a title that reads less like content marketing and more like a plea: "Do GEO/AI Visibility Platforms/Tools Matter? Help Me Find Out." Forrester spent years inside Bing and later Yext, which means he's watched more than one hype cycle promise a measurement layer that turned out to be mostly theater. He isn't a junior analyst chasing a trend piece. He built an entire, deliberately skeptic-framed, crowdsourced industry survey because he does not currently know the answer — and he wanted the people actually using these tools to tell him, not the vendors selling them.
Sit with that for a second before you get to the results, which won't land until August anyway. A person with this much institutional memory in search didn't run a case study or a customer testimonial roundup. He ran an open industry survey explicitly designed to surface the skeptical view — whether the entire AI visibility platform category is solving a real problem or manufacturing one. You don't build that kind of survey about a category you already trust. You build it about a category you suspect is running ahead of its own evidence.
That's the headline, months before any results post. Not "GEO platforms proven" or "GEO platforms debunked." The headline is that the question needed to be asked this formally, by someone this senior, in public, with a request for practitioners to send him receipts. If the category's value were self-evident, nobody credible would need to go looking for it.
Why an AI Visibility Platform's Score Can't Be Audited
Here's the mechanical issue underneath Forrester's question, and it's one we've written about before because we keep running into it with new clients. Most AI visibility platforms sell you a single composite number. Call it a visibility score, a presence index, a share-of-voice metric — the label changes, the structure doesn't. Somewhere in a vendor's backend, mention counts across several AI engines get weighted, normalized, and blended into one figure that's meant to represent your entire footprint in AI answers. It's clean. It's screenshot-friendly. It's also, in most implementations, not something you can independently check.
A dashboard is not proof. It's a picture someone sold you. When that score climbs 12 points month over month, you have no reliable way to know why. Maybe you're genuinely winning more citations. Maybe the vendor expanded its prompt sample. Maybe they swapped which AI engines get weighted this quarter, or changed how they define a "mention" versus a "citation" versus a "reference." We went deep on this exact failure mode in why your mention-rate dashboard is lying to you — the short version is that a blended score hides more signal than it reveals, and it hides it by design, because a single number is what fits on a slide.
None of this makes the vendors dishonest. It makes the metric unfalsifiable, which is a different and in some ways more dangerous problem. A dishonest number gets caught eventually. An unfalsifiable one just sits there, quietly steering budget, because nobody on the customer side has the raw inputs to check it against reality. You paid for the platform. You do not own the methodology.
What a Blended Score Can and Can't Tell You
This isn't a data table pulled from a study — there isn't one yet, Forrester's results aren't published until August. Think of it instead as a framework for a conversation you should be having internally right now, before you renew or buy anything: for each question a stakeholder is actually going to ask you, can the vendor's blended score answer it, and can your own engine-level instrumentation answer it?
| QUESTION A STAKEHOLDER WILL ACTUALLY ASK | VENDOR'S BLENDED SCORE | YOUR OWN ENGINE-SPECIFIC INSTRUMENTATION |
|---|---|---|
| Are we cited more on ChatGPT specifically, or Perplexity, or both? | No — engines are averaged together | Yes — tracked separately per engine |
| Did we gain real citations, or did the vendor change its sample? | No — methodology is proprietary | Yes — you control the query set and can rerun it |
| Where do we rank among the sources an engine actually cites? | Rarely — most scores skip rank entirely | Yes — citation rank is a directly loggable number |
| How often does an engine retrieve our content at all, cited or not? | No — retrieval and citation get conflated | Yes — retrieval frequency is a separate, trackable signal |
| Can we defend this number to a CFO who asks how it's calculated? | Difficult — the formula is a trade secret | Yes — every input is one you logged yourself |
Read down that right-hand column and the pattern is obvious: none of it requires a platform. It requires discipline. Different AI engines run on genuinely different retrieval mechanics — we've mapped how AI visibility tracking swings differently by market share depending on which engine you're measuring — which is exactly why averaging them into one score throws away the information that actually tells you where to spend effort.
What to Build Instead of an AI Visibility Platform
So skip the composite score and track the three things underneath it. None of these require enterprise software. They require someone willing to run the same prompts on a schedule and log what comes back, the same unglamorous way analytics has always worked.
We built out the full mechanics of this in measuring GEO: grounding, citation, and market share as separate metrics, and the per-engine investment logic in our generative engine optimization framework, because the instrumentation only pays off if you route budget differently once you have it. A platform that tells you "visibility is up" gives you nothing to act on. Knowing that ChatGPT citation rank slipped for your category's comparison queries while Perplexity retrieval held steady tells you exactly which page to fix first.
If your own reporting stack is thin on this, that's exactly the gap our reporting and analytics work exists to close — not by selling you a new composite score, but by wiring up the engine-specific numbers underneath it so you can defend them to anyone who asks.
You're Not Behind for Skipping the Platform
Every category running this hot produces a version of this anxiety. If everyone else has a platform and you don't, it feels like you're behind. Resist that feeling for a minute, because it's measuring the wrong thing. The teams that get burned aren't the ones who waited to buy a GEO platform. They're the ones who have no real numbers at all — vendor-supplied or homegrown — and are steering a real budget on vibes, screenshots, and whichever citation someone happened to notice in a chat window last week.
“You're not behind for not having bought an AI visibility platform yet. You're behind if you have no real numbers at all — vendor or homegrown.”
Forrester's survey will publish results in August, and they'll be worth reading closely when they do. But you don't need to wait on them to change how you operate this month. Whether the eventual data says the category earns its price tag or mostly doesn't, the fix underneath it is the same either way: stop trusting a blended score you can't reconstruct, and start logging the three numbers that actually explain what's happening — engine by engine, citation by citation, retrieval by retrieval. That's not a fallback plan until a real platform comes along. For a lot of teams, it's the more durable plan, full stop.
A vendor score can make you feel measured without making you informed. The reframe worth carrying out of this piece isn't "platforms bad." It's that measurement is a discipline you can own, not a subscription you rent — and the fact that someone as seasoned as Forrester had to go crowdsource proof of value is your permission slip to ask the same skeptical question about whatever's currently sitting on your own dashboard.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.