Every AI visibility platform on the market reports the same kind of number: a mention rate, a citation count, a share-of-voice line moving up or down. None of them answer the question that actually determines whether a GEO budget is well spent, which is where the citations that make up that number are structurally coming from, and whether that structure is stable enough to plan against. We cross-referenced five independently published 2026 datasets, none of them ours, to answer that question directly, and built a new scoring framework out of what they show together that none of them shows alone.
Methodology: What No AI Visibility Platform Measures
None of the five sources behind this piece were produced by Something Inc., and we want that stated before anything else. What follows is a synthesis: five independently published, dated 2026 datasets and reports, each measuring a different slice of how AI engines select and surface sources, read together to answer a question none of them individually set out to answer. Is AI search source selection concentrating around a durable, structural set of winners, or is it still an open, evenly distributed field? And if it is concentrating, how exposed is any given brand or category to that concentration shifting under it without warning?
The first source is Ahrefs founder Tim Soulo's summary of a 1-billion-data-point analysis across 14 separate studies, first surfaced in a June 25, 2026 LinkedIn post and subsequently written up by bloggersideas.com. It measures what AI engines actually cite once they answer a query: which source types dominate, and how much overlap exists between engines and even between two products from the same company.
The second source is Duane Forrester's essay "Why Part of Your AI Authority Takes Years, Not Campaigns," published August 2, 2026. It doesn't measure citations directly. It explains the mechanism underneath a specific kind of citation, what a model says about a brand purely from its trained weights, with no live retrieval involved, and why that mechanism moves on a multi-year timeline no campaign can compress.
The third source is Cloudflare Radar referral-traffic data, aggregated and published by TechnologyChecker.io on a page last updated August 1, 2026. It measures the one number most executives actually see on a dashboard: what share of tracked referral traffic across the open web comes from Google versus everyone else, month over month.
The fourth source is a July 23, 2026 Search Engine Roundtable writeup by Barry Schwartz of a Wall Street Journal report (resurfaced August 4, 2026), naming USA Today, Reuters, Politico, People Inc., The Economist, and Time as publishers weighing blocking Google's crawler outright over lost traffic tied to AI Overviews. It's practitioner behavior, not a metric, but it's practitioner behavior that only makes sense if something the aggregate number in source three doesn't show is actually happening underneath it.
The fifth source is another Schwartz report at Search Engine Roundtable, "Meta/Facebook Crawling Web & May Be Building Search Engine," published August 10, 2026. This is early-stage, single-source evidence, not a launched product or a confirmed roadmap, and we're flagging it that way explicitly rather than treating it as settled. It matters here for a narrower reason: it's a live example of how fast the competitive set an entire framework is built around can change.
Five sources, five different measurement methods and evidentiary weights, from citation-composition analysis to mechanism-level explanation to aggregate traffic data to publisher behavior to early competitive signal. None of them, alone, supports the conclusion this piece reaches. Read together, against each other, they do.
Finding One: The Winners Are Structural, Not Random
Start with what actually gets cited. Soulo's summary of the 14-study, 1-billion-data-point analysis breaks citation sources down by type, and the breakdown is not a long tail. Wikipedia accounts for 29.7% of citations. Brand and company homepages account for 23.8%. App stores account for 6.6%. Add those three together and a single domain type, an encyclopedia, plus one structural category, the brand's own homepage, plus one more, app-store listings, account for well over half of everything an AI engine chooses to cite across the studies measured.
That is not a random distribution across millions of individual domains. It's concentration around a small number of source types that share a structural property: each one is a canonical, single-answer location for a given entity. There is one Wikipedia page per topic. There is one homepage per brand. There is one app-store listing per product. AI engines appear to default to canonical single sources over the fragmented, many-competing-pages model that classic organic search rewarded for two decades. This is exactly the kind of structural pattern an ai citation tracking program needs to be built around, not a coincidence to note in passing.
Share of AI citations by source type, across 14 studies and 1B+ data points (Ahrefs / Tim Soulo, Jun 2026)
The same analysis found that 67% of ChatGPT's top citations come from sources marketers don't directly control. Read next to the source-type breakdown above, that's not a contradiction, it's the same finding from a different angle. Wikipedia isn't owned by any brand. App-store listings are partially controlled, at best, by store algorithms and review aggregation a brand can influence but not dictate. Even a brand homepage, the one category on that list a company fully owns, is competing for a single canonical slot per entity rather than for one ranking position among ten. Owning your homepage doesn't guarantee it wins that slot against a stronger third-party page covering the same entity.
This is the first half of what makes source concentration a structural phenomenon rather than a random one: winners cluster around canonical-source categories, and a majority of the highest-value real estate in those categories sits outside direct brand control. A generative engine optimization program that spends its budget purely on owned content is optimizing for the minority of the pie, not the majority.
Finding Two: Google Doesn't Agree With Itself
Here is the finding in the Ahrefs analysis that should unsettle anyone running a single, blended GEO strategy across engines: AI Mode and AI Overviews, both Google products, reach the same conclusion 86% of the time, but cite different sources to get there 86.3% of the time. Only 13.7% of citations overlap between the two. Two products from the same company, trained substantially on the same underlying infrastructure and index, agree almost entirely on what to say and almost entirely disagree on where to say it came from.
That single data point breaks an assumption most enterprise GEO planning still runs on implicitly: that if you win Google's AI-powered surfaces once, you've won Google. You haven't. You've won one retrieval and ranking pass out of at least two Google runs in parallel, with an 86.3% chance the other one drew from an almost entirely different citation pool to reach the same answer. This is the empirical case for treating per-engine investment as its own discipline rather than a single blended plan, and it now applies within a single company's own product suite, not just across ChatGPT, Claude, Perplexity, and Copilot.
| COMPARISON | WHAT MATCHES | WHAT DOESN'T |
|---|---|---|
| AI Mode vs. AI Overviews (same company) | 86% conclusion agreement | 86.3% citation-source disagreement (13.7% overlap) |
| Google aggregate share, May vs. Jul 2026 | 87.63% → 88.08% referral share (essentially flat) | Underlying source composition and publisher sentiment shifting hard underneath it |
The mechanism explains part of why this happens, and Forrester's essay is where that mechanism gets named directly. A meaningful share of what a model says about any given entity comes from what he calls parametric standing, the model's baked-in sense of a brand or topic drawn from its trained weights, independent of whatever live retrieval pulls in at query time. Parametric standing accumulates slowly, from years of independent third-party mentions absorbed into a training corpus, and it cannot be built quickly through a campaign. If AI Mode and AI Overviews weight parametric standing against live retrieval even slightly differently in how each product assembles its answer, two runs against overlapping knowledge can land on the same conclusion through almost entirely different citation paths. This is a structural property of how these systems are built, not a bug either product will patch away, and this dynamic is exactly why our engines disagree on sources even when they land on the same answer.
Forrester's underlying research on the training corpus itself sharpens why parametric standing is so slow to move. The C4 training corpus is dominated by content written between 2011 and 2019, which makes up 92% of it. A single domain, no matter how authoritative, represents less than 0.05% of that corpus. Research by Allen-Zhu and Li shows reliable fact recall from a model requires sufficiently varied phrasing across many independent mentions during pretraining, not one well-optimized page. Research by Mallen et al. shows that scaling model size helps a model get more confident about popular, already-well-covered entities, but leaves long-tail entities roughly where they started. Put together: if your brand isn't already well-represented across a wide, independently authored slice of a decade-plus of web content, no amount of model scaling alone closes that gap, and no single quarter of content production closes it either.
Finding Three: The Aggregate Number Is Lying by Omission
Now the part of this synthesis that ties directly to what executives actually watch. TechnologyChecker.io's Cloudflare Radar data, on a page last updated August 1, 2026, shows Google sending 87.63% of tracked referral traffic in May 2026 and 88.08% in July 2026. That's not a decline. It's not even a meaningful move. By the one number most dashboards surface to leadership, Google's dominance of referral traffic looks completely stable across the exact window this entire framework is built on.
Read that number next to Schwartz's July 23, 2026 report of the Wall Street Journal story, and the two don't fit together comfortably. USA Today, Reuters, Politico, People Inc., The Economist, and Time are, per that reporting, weighing blocking Google's crawler entirely because of traffic loss tied to AI Overviews. Publishers of that size don't consider blocking the largest referral source on the internet over a shift that doesn't exist. They're reacting to something real, happening inside the composition of how Google sends (or doesn't send) traffic, that the aggregate 87.63%-to-88.08% share figure never captures, because that figure only measures total share against other referral sources, not the underlying mix of how much of that Google traffic converts to an actual visit versus getting answered and absorbed on the results page itself.
Google's share of tracked referral traffic, May vs. Jul 2026 (Cloudflare Radar via TechnologyChecker.io, updated Aug 1, 2026)
This is the clearest evidence in this synthesis for a specific, underappreciated risk: the metric your organization is watching can look completely stable while the thing it's supposed to represent is moving underneath it. An aggregate referral-share number aggregates away exactly the kind of shift that individual publishers, close enough to their own traffic to feel it directly, are already restructuring their crawler policy over. If the largest, best-resourced publishers in the market are making infrastructure decisions based on a shift a stable top-line number doesn't show, a marketing team relying on that same top-line number to greenlight or table a GEO investment is working from an incomplete picture, and this is a big part of why we've argued elsewhere that a single mention-rate dashboard can actively mislead the team reading it.
It also reframes what an ai search market share number is actually useful for. A share-of-referral or share-of-citation figure answers how big is my slice, which is a fine question for a board slide. It does not answer whether the composition of that slice, and your exposure to it changing, is stable enough to plan next year's budget against, which is the question that actually determines whether a GEO program is under- or over-invested relative to real risk.
Finding Four: A New Entrant Could Reshuffle All of It
Everything above describes a snapshot of a competitive field with, effectively, one dominant search infrastructure provider (Google, across its AI Mode and AI Overviews products) plus a set of independent AI-native challengers (ChatGPT, Claude, Perplexity, Copilot). Schwartz's August 10, 2026 report at Search Engine Roundtable adds a variable none of the other four sources account for: early-stage, single-source evidence that Meta is crawling the web at scale and may be building its own search infrastructure.
We're stating the evidentiary weight of this plainly because it matters to how much of a category's planning should lean on it: this is one outlet's reporting, based on observed crawler behavior and pattern-matching against Meta's public statements, not a confirmed product, an announced launch date, or a company-verified roadmap. Treat it as a signal worth tracking, not a fact to plan a budget around today.
What it demonstrates regardless of whether Meta ships anything is the fragility of any framework, including this one, that treats today's competitive set as fixed. Every finding above, the source-type concentration, the intra-Google citation disagreement, the gap between aggregate stability and ground-level publisher behavior, was built by measuring the current field: Google's two AI products, plus ChatGPT, Claude, Perplexity, and Copilot. A new entrant with Meta's scale of first-party user behavior data and existing crawler infrastructure wouldn't just add a sixth engine to track. It would very plausibly cite differently than any of the five above, at the same time that a brand with billions of users, an app-store presence, and homepage-equivalent profile pages across its own properties could itself become one of the structural winner categories described in Finding One. That's a second-order effect worth naming even at this early a stage: a Meta search product wouldn't only compete for query share, it could become a new type of canonical, structurally-favored source in its own right.
The Framework: Source Concentration Exposure
Four findings, read separately, are four interesting data points. Read together, they describe a specific kind of risk that no single AI visibility platform metric currently scores: how exposed a brand or category is to structural concentration around a narrow set of source types, to disagreement between engines (and within a single engine's own product family) about which sources in that narrow set actually get cited, to a slow-moving training-corpus layer that a campaign can't accelerate, and to an aggregate metric that can stay flat while the composition underneath it moves. We're calling that composite risk source concentration exposure, and scoring it requires four inputs, each pulled directly from one of the findings above.
The first input is source-type concentration: how much of the citation volume in a given category clusters around a small number of canonical source types (encyclopedic references, brand homepages, app-store or directory listings, a small set of trusted trade publications) versus spreading across a genuinely long tail of independent domains. Categories with high concentration, closer to the 30%-plus Wikipedia share Finding One documents, carry more exposure, because winning or losing access to that narrow set of canonical slots swings a disproportionate share of total citation volume.
The second input is cross-engine agreement: how consistently ChatGPT, Claude, Perplexity, Copilot, and Google's own AI Mode and AI Overviews draw from the same sources when answering comparable queries in a category. Finding Two's 13.7% overlap figure, measured within a single company's own product suite, sets a useful low-agreement benchmark. A category scoring near that level needs a genuinely separate strategy per engine. A category with high cross-engine agreement can run a more unified plan.
The third input is training-corpus dependency: how much of a category's typical AI answer draws on parametric standing (the slow-accumulating, trained-weights layer Forrester describes) versus live retrieval from current web content. Categories where AI answers lean heavily on long-established brand reputation, the kind built from a decade-plus of independent, varied third-party mentions, carry a different kind of exposure than categories where answers are assembled fresh from recent content each time. The former is a multi-year asset that under-invested brands can't buy their way into quickly; the latter rewards active content and structure work far faster.
The fourth input is aggregate-metric exposure: how much a brand or category's GEO strategy currently relies on a single top-line stability number (a referral-share percentage, a blended mention rate, a citation-count trendline) as its primary signal, without a parallel check on composition. Finding Three shows this exposure can be severe even when the top-line number itself is genuinely accurate; the risk isn't that the number is wrong, it's that stability at that level of aggregation is compatible with real disruption one layer down.
A category or brand that scores high on all four axes, tightly concentrated sources, low cross-engine agreement, heavy training-corpus dependency, and heavy reliance on an aggregate number, is the highest-exposure profile this framework describes: hard to win quickly, hard to plan with one blended strategy, slow to show movement in a reporting period, and prone to a top-line metric that looks fine right up until it doesn't. That combination maps closely onto what Kevin Indig's category-ownership research already found about how few AI-search categories have a stable owner at all; a category that's mostly unowned is, almost by definition, one where the structural winners described in Finding One haven't been fully claimed yet, which raises the stakes on getting the concentration read right early rather than after a competitor locks in the canonical slot.
Scoring Your Category Beyond a Standard AI Visibility Platform
None of the four inputs above require a proprietary dataset to estimate directionally. Source-type concentration can be approximated by pulling the top-cited domains for your category's core query set across engines and checking how many distinct source types (not just distinct domains) show up in the top results; a category dominated by two or three types is high-concentration regardless of how many individual domains sit inside those types. Cross-engine agreement can be approximated the same way Ahrefs approximated the AI Mode versus AI Overviews gap: run the same query set through multiple engines, or through a single provider's multiple products where available, and measure citation overlap directly rather than assuming it. Training-corpus dependency is the hardest to measure precisely without direct API access to a model's internals, but it can be approximated by comparing what a model says about your brand with no browsing or retrieval enabled against what it says with retrieval enabled; a large gap suggests heavy reliance on live sources, while a small gap suggests the model already has a strong parametric position independent of what's currently live on the web. Aggregate-metric exposure is the simplest to check of the four: ask whether your current GEO reporting has a single number an executive would point to as the indicator, and whether anyone on the team could tell you, without pulling a fresh report, whether the composition behind that number shifted in the last quarter.
This is also where an honest audit of GEO readiness earns its keep over a dashboard subscription: a platform that reports mention rate or citation count is answering a different question than source concentration exposure asks, and the two are not substitutes for each other. A category can show a stable or even rising citation count on a standard AI visibility platform while scoring dangerously high on every one of the four exposure axes above, concentrated sources, low cross-engine agreement, heavy corpus dependency, and a leadership team watching one number that hides all three.
None of the five sources behind this framework, taken alone, would have supported building it. The Ahrefs citation-composition data shows concentration exists but says nothing about whether that concentration is stable across engines. Forrester's mechanism explains why some of that concentration is slow to shift but doesn't quantify current citation share. The Cloudflare Radar data shows aggregate stability but can't, by design, show composition. The WSJ publisher story shows real disruption but is anecdotal, not a metric. The Meta report is a single early signal about tomorrow's competitive field, not a measurement of today's. It's the cross-referencing itself, treating five independent measurements of five different layers of the same system as inputs to one model, that produces a framework none of the five original authors set out to build.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.