Picture the call. A client opens the deck, points at one line, and asks why their AI visibility score dropped twelve points this month. Nobody on the call can answer with any confidence. Not because the team is bad at the job, but because the tool that produced the number doesn't show its work, and neither does the report built on top of it. That call happens somewhere every single week right now, and it's the best evidence I know that GEO reporting has an honesty problem before it has a measurement problem.
The uncomfortable part is that this isn't a hunch or a vibe picked up from one bad client meeting. In August, Duane Forrester, who spent years running search products at Bing and Yext before he started writing Duane Forrester Decodes, surveyed 163 practitioners running GEO programs about how much they trust the platforms measuring their AI visibility. The headline finding wasn't a scandal. It was quieter than that, and worse: a meaningful share of the people running these programs day to day don't fully trust the numbers their own tools hand them. Not a competitor's tool. Their own.
A trust problem 163 practitioners just put a number on
Sit with that for a second. These aren't casual users clicking around a free trial for an afternoon. These are people whose job is to run GEO programs, defend budgets in front of skeptical finance teams, and explain a swing to a client every single month. If the people closest to the tooling are hedging on what it tells them, the client three steps removed from the methodology has no real chance of trusting it either. A dashboard the operator doesn't fully believe isn't a measurement problem you fix with a nicer chart. It's a credibility problem, and credibility problems compound instead of averaging out over time.
“A number without a method isn't data. It's a rumor with a decimal point.”
Why AI visibility reporting collapses into one score
Here's the part nobody selling GEO tools wants to say out loud: a single number is easy to sell. 'Your visibility score is 62' fits in a headline, a slide, a renewal email. It travels well through an org chart, from analyst to director to the client's CMO, without losing anything on the way up. A methodology section does not travel that well. It has caveats. It has footnotes. It forces a buyer to think for one extra second before nodding, and a vendor market optimized for smooth nods will not volunteer that friction on its own. So the market optimized for the thing that sells, not the thing that's fully true, and most GEO dashboards now hand clients one blended figure standing in for a dozen decisions nobody shows them.
We've made pieces of this case before, from a few different angles, because the pattern keeps showing up in the data we look at. We've asked whether AI visibility tracking is a vanity metric in disguise, and the honest answer is that it depends entirely on the method behind the number, which is usually the one thing the report doesn't show. We've written about AI search market share numbers that don't agree with each other even when two vendors claim to be measuring the same engines in the same month. And we've shown how forced-search queries can quietly skew what a visibility tool reports, producing a number that looks perfectly stable while the underlying behavior it's supposed to represent is anything but. Different angle, same finding every time: the blended number is the least trustworthy part of the report, and it's the only part most clients ever actually see.
None of this makes the tools useless, and it would be a mistake to read it that way. It makes them incomplete as delivered. A GEO platform that logs mention rate across a defined set of engines, on a defined date, from a defined prompt set, is doing real, checkable work. The problem starts the moment that work gets compressed into a single number for the client deck and everything that made it checkable gets left in the tool's back end, visible only to the analyst who happened to build the query set.
The GEO reporting methodology label
Nutrition labels exist because 'contains sugar' used to be entirely a marketing decision, not a disclosure. Now it's a format. Every package carries the same handful of facts in the same order, in roughly the same place, and a shopper who wants to check can check in five seconds without reading a research paper on the ingredient. GEO reporting is roughly where food labeling was before that standard existed: sellers control the one number buyers see, and buyers have no standard, repeatable way to ask what's underneath it. That's fixable, and it doesn't require a new tool, a new vendor, or a committee to argue about it for a quarter. It requires attaching a short, consistent set of methodology facts to every AI visibility number a team reports, every single time, with no exceptions for the months the number looks good.
| FIELD | WHAT IT TELLS THE READER | WHY IT'S NON-NEGOTIABLE |
|---|---|---|
| Engines queried | Which AI systems the number covers: ChatGPT, Perplexity, AI Mode, Claude, or some subset of them | A 'visibility score' that's actually 80% one engine reads very differently once the client knows that |
| Forced or natural search | Whether the tool made the assistant search on every single query, or let the assistant decide on its own | Forced search inflates citation opportunities the assistant might never have taken naturally, which inflates the score with it |
| Sample size | How many prompts the number is actually built from | Twenty prompts and two thousand prompts should never be presented on a slide with equal confidence |
| Date window | The exact start and end date the data was pulled | AI answers shift week to week; a number with no date attached is already stale before the meeting starts |
| Logged-in or logged-out | Whether the queries ran through an account with history, or a clean, anonymous session | Personalization changes what gets cited; a logged-in run isn't a fair comparison to a logged-out one |
| Query set source | Who wrote the prompts, and how they were chosen | A prompt list built by the vendor to flatter itself is a different instrument than one built from real buyer language |
What the label actually contains
None of this is exotic, and none of it requires new engineering. It's the same handful of facts a competent analyst already tracks privately to sanity-check their own numbers before a client call. The only real change we're proposing is making that private checklist public, attached to the report itself, in the same five or six lines, every time. Teams running generative engine optimization programs for B2B SaaS clients in particular tend to have stakeholders who ask sharp methodology questions anyway, often engineers or data people sitting on the buying committee. Hand them the label up front and the follow-up meeting turns into agreement instead of an interrogation nobody enjoys.
What changes when the method is visible
The instinct against this is understandable. Nobody wants to hand a client a longer report, and nobody enjoys explaining that last quarter's number was built on a smaller sample than this quarter's. But the alternative is what Forrester's 163 practitioners are already living with: a slow drain of confidence in numbers nobody can fully vouch for, including the people whose actual job is to vouch for them. A visible method doesn't just protect the client from a bad read. It protects the team reporting the number, because a swing that looks alarming in isolation often reads as expected once the sample size and date window sit right next to it on the same page.
There's a quieter benefit too, one that only shows up after a few reporting cycles. Once a client has seen the label a handful of times, they stop asking you to defend the number and start asking better questions about the program itself: which query set should we expand next, which engine is worth chasing harder, whether the sample needs to grow before a board update. That's a genuinely different conversation than defending a score nobody can see inside. It moves the relationship from policing the measurement to actually running the program, which is the conversation every GEO team wants to be having in the first place.
We built our own reporting and analytics work around this same instinct well before Forrester's survey gave it a name: don't hand a stakeholder a number you wouldn't want audited on the spot. Look at how a real engagement gets reported once the label is standard practice, and you'll notice the same pattern repeated: the win gets stated plainly, and the method that produced it sits one click away, not buried in an appendix nobody opens. That's not extra transparency for its own sake, and it's not busywork added to satisfy a compliance checklist somewhere. It's the difference between a client who trusts the next number you show them and one who's quietly building a spreadsheet of their own, on the side, to check your math against.
So here's the permission, if you need it: you don't have to wait for the GEO tooling market to fix its trust problem before you fix your side of it. You don't own the model. You don't own the crawler. You do own the report that goes out under your name, and that report can carry a real method starting with the next one you send.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.