Ask around your own marketing team, quietly, whether anyone can actually prove the AI visibility work is paying off this quarter. Not "we think it's helping." Not "citations feel like they're up lately." An actual number, defensible in a budget meeting, tying spend to outcome. If the room goes a little quiet, you're not behind. You're in the 45%.
The question everyone's afraid to ask out loud
Here's the version of the question nobody wants to say in the room: we've been talking about AI visibility for over a year now, we've shipped llms.txt files and schema and comparison content, and if someone asked you right now to show the ROI on all of it, could you? Not the citation count. Not "ChatGPT mentioned us in a screenshot someone found." An actual causal line from the work to a business outcome.
Semrush put a number on how many marketing leaders can answer that honestly, and it explains a lot about why the question keeps getting dodged instead of answered.
It's worth naming why this particular gap feels so much worse than the usual "we should track this better" admission every marketing team makes about something. Traditional SEO reporting, whatever its flaws, has two decades of tooling behind it. Rank trackers, click-through data, attribution models, all imperfect, all argued over, but all fundamentally capable of answering the basic question of whether a page is performing. AI visibility doesn't have that inheritance yet. There's no equivalent of Google Search Console for ChatGPT. There's no universal, vendor-neutral way to see exactly which of your pages an engine pulled from, how often, or what happened after. Teams are being asked to justify budget for a channel that doesn't yet offer the basic instrumentation the previous twenty years of digital marketing trained everyone to expect.
That's not an excuse for staying in the 45%. It's context for why so many otherwise sophisticated marketing organizations are sitting there anyway.
126 million prompts and a 45% blind spot
The 2026 AI Visibility Index scaled up from an initial 2,500-prompt pilot to 126 million U.S. AI search prompts, analyzed from January through April 2026 across ChatGPT, Gemini, Google AI Mode, and Google AI Overviews. Buried inside that dataset is a much less flattering number than the headline scale suggests: 45% of marketing leaders say they cannot accurately measure their brand's visibility in AI-generated answers. Only 9% say they have tools that track all the metrics that actually matter.
Sit with the 9% for a second, because it's honestly the more damning number of the two. Almost half the entire market admits it can't measure this well at all. Fewer than 1 in 10 say they have full coverage. That means even inside the 55% who claim they can measure it, most are still working from partial visibility, one platform tracked reasonably well, three tracked poorly or not tracked at all, and reporting that partial picture upward as though it were the whole, complete story.
The 81% versus 36% gap that actually explains it
The most useful number in the whole study isn't the 45% who can't measure. It's the split between organizations that treat SEO and AI visibility as one integrated workflow and organizations that manage them as two separate lanes. Among the integrated group, 81% reported increased traffic or leads from AI platforms. Among the siloed group, only 36% reported the same.
| WORKFLOW STRUCTURE | REPORTED TRAFFIC/LEAD INCREASE FROM AI PLATFORMS |
|---|---|
| Fully integrated SEO + AI visibility | 81% |
| Managed as separate workstreams | 36% |
That 45-point gap is too large to be a measurement artifact. It's more likely a genuine outcome gap, and the mechanism isn't mysterious once you think about how AI engines actually source answers: the same technical and content signals that drive organic rankings, schema, extractable structure, authoritative corroboration, are heavily overlapping with the signals that drive AI citation. Teams that run these as one program get to reuse the same data, the same content pipeline, and the same measurement infrastructure across both. Teams running them separately are often duplicating work, missing the overlap entirely, or worse, watching one team's SEO change accidentally undo the other team's GEO progress because nobody's coordinating.
If your team is one of the 64% still running these as separate lanes, that gap is the business case for merging them, and it's a cleaner argument than most of the abstract "AI is the future" pitches that tend to get this kind of initiative funded in the first place. This is a measured outcome difference, not a prediction.
It also reframes what "AI visibility" spend is actually buying. Treated as its own separate line item, it looks like a speculative bet on a still-forming channel, easy to defer, easy to cut when budgets tighten. Treated as one workflow with SEO, it looks like what it actually is: the same technical foundation and content investment doing double duty across two increasingly overlapping surfaces. Framing it the second way isn't just more accurate, it's a materially easier number to defend in the room where the budget actually gets decided, because it stops competing against SEO for the same dollars and starts compounding with it instead.
Why ChatGPT and Gemini make this even harder
Part of why a single "AI visibility score" is so hard to build honestly is structural, not just a tooling gap. Semrush's data shows ChatGPT citing an average of 15 sources per response, while Gemini cites around 3. A brand that shows up reliably across ChatGPT's wider citation pool can look strong on a blended score while being nearly invisible on Gemini specifically, and a blended number hides exactly that kind of gap.
This lines up with what we found looking at a separate study that asked three engines the same 30 questions and got 104 different cited sources back: engines don't agree with each other nearly as much as a single visibility dashboard implies when it reports one composite number. Measuring "AI visibility" as though it were one metric is part of why 45% of leaders can't get a straight answer out of their own reporting. The honest version of this measurement requires per-engine breakdowns, not a single score that averages away the platforms where you're actually weak.
There's a practical trap hiding in that averaging effect, too. A composite score that's rising month over month feels like good news, and it might be, but it can just as easily mean you're getting stronger on the engine where you were already strong while quietly losing ground somewhere you never check. Without a per-engine breakdown, those two very different stories produce the exact same upward-trending line on the one dashboard most teams actually look at, and only one of them is a story worth reporting to leadership as a win.
What to do before your next budget conversation
You don't need to solve all of this before your next budget review, but you do need an honest answer to two questions, because "I'm not sure" is the answer that gets AI visibility spend cut first when budgets tighten.
The order those three questions get answered in matters more than it might seem. Fixing the dashboard before fixing the team structure just produces a more precise version of the same siloed, underperforming setup, a beautifully broken-out per-engine chart that two disconnected teams still can't act on coherently. Merging the workflow first, even with imperfect measurement in place at the start, is closer to what the 81%-versus-36% data actually describes: the organizational structure predicted the outcome more strongly than the sophistication of the reporting did.
This is also where reporting and analytics work earns its keep in a way that's easy to underrate until you're the one standing in a budget meeting without an answer. A dashboard that ties per-engine citation rate to organic and AI-cited pipeline in one view is the difference between defending this spend with a number and defending it with a feeling. For B2B SaaS teams where AI-assisted research increasingly happens before a prospect ever fills out a form, that's not a nice-to-have reporting upgrade. It's the only way to know, with any real confidence, whether the work is actually working.
None of this requires waiting for the industry to build a universal AI-search analog to Search Console before you start. It requires picking the handful of prompts that matter most to your pipeline, checking who gets cited for each one across every engine your buyers actually use, and building the habit of reporting that per-engine, tied to pipeline, on a fixed cadence. That's a genuinely smaller, more achievable project than it sounds at first, and it's the difference between staying in the 45% by default and being able to answer the question honestly and specifically the next time someone in the room actually asks it.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.