Something Inc.Schedule a free consultation
GEO

Google Says No Site Gets Special Treatment in AI Search. The Crawl Data Disagrees.

Google told The Verge Reddit gets no special preference in its rankings or AI features. Fair enough. But that was never really the question worth asking about ChatGPT citations.

JBJosh BernsteinManaging Partner · AUG 6, 2026 · 10 MIN READ

I've now read the same Google statement three times this week, forwarded to me by three different clients, each with a version of the same question attached: is this true? Does Reddit really get no boost? And every time, I've had to explain that the statement is probably accurate and also not the thing anyone should be arguing about.

KEY TAKEAWAYGoogle denying a preference rule is not the same as denying a structural advantage. The domains that dominate ChatGPT citations and AI Overviews — Reddit, Wikipedia, YouTube, Google's own properties — likely get there through how AI training data mirrors the web's existing authority graph, not through a switch anyone flipped. That's a harder problem to out-argue and a more useful one to build against.

Here's what happened. On August 5, a Google communications rep named Jennifer Kutz told The Verge that Reddit "gets no special preference" in Google Search or in Google's AI features. She said the company's systems are built to tell helpful user-generated content apart from low-quality content, and added the standard line: "There will always be bad actors trying to game the system...we have strong protections against manipulation across Search, including our AI features." Search Engine Roundtable covered it the same day, and if you work in SEO or GEO, your feed probably surfaced it within the hour.

The statement that answered a different question

Some context first, because the statement didn't come out of nowhere. Reddit has spent months publicly grumbling about falling Google traffic, and its preferred explanation is AI Overviews eating clicks before users ever reach the thread. Google's counter-framing, offered here and elsewhere, is that any decline is more likely explained by core algorithm updates and spam-related changes than by anything Reddit-specific. Two large platforms, both with something to protect, both offering the explanation that's kindest to their own roadmap. Normal Tuesday in search.

Set the politics aside and just take the claim at face value: no hard-coded rule that says "give Reddit a boost." I believe that. Nobody credible I know in this industry has ever claimed Google keeps a literal allowlist with Reddit's name typed into it. That was never the sharp version of the argument, which makes the denial a little bit like reassuring someone their smoke detector isn't haunted when what they actually asked was why the kitchen keeps catching fire.

The real question — the one buried under every panicked Slack message about ChatGPT citations this year — isn't whether an engineer wrote a preference rule. It's why the same small set of domains keeps showing up at the top of citation data across completely different companies, different models, and different retrieval systems, none of which coordinate with each other. That's a pattern too consistent to be coincidence and too structural to be a favor.

What the Common Crawl data means for AI citations

This is where a piece of research from earlier this year gets useful, and it's the actual reason I wanted to write this instead of just retweeting the Verge story with a shrug. Metehan.ai published an analysis in January arguing that a genuinely underexamined signal in AI search citations is Common Crawl's own domain-authority data — specifically Harmonic Centrality rank and PageRank, both computed over Common Crawl's WebGraph. The logic is straightforward once you say it out loud: most large language models are trained heavily on filtered Common Crawl data. Sixty-four percent of the 47 LLMs the piece analyzed, spanning 2019 through 2023, used filtered Common Crawl data in training. GPT-3 reportedly derived more than 80% of its training tokens from filtered Common Crawl sources.

If a huge share of what a model learned about the world came pre-weighted by which domains the web already treats as authoritative, then the model's instinct about which sources sound trustworthy didn't come from nowhere. It came from the same graph that made Reddit, Wikipedia, and Google-owned properties central to the open web a decade before anyone was optimizing for ChatGPT citations. That's a very different claim than "Google boosts Reddit." It's closer to: the raw material every model was trained on already had Reddit near the center, so of course models keep reaching for it.

The citation-share numbers back this up loudly. A Semrush analysis of more than 150,000 citations found Reddit accounted for 40.1% of them, with Wikipedia at 26.3%, Google at 23%, and YouTube at 23% — note these overlap rather than sum to 100%, since a single answer can cite more than one of them. Separately, Profound's analysis of 680 million citations found Wikipedia leading ChatGPT citations at 7.8% and Reddit leading Perplexity citations at 6.6%. Different methodology, different scale, same handful of names at the top.

Reddit40.1%
Wikipedia26.3%
Google23%
YouTube23%

Share of 150,000+ analyzed AI citations, by domain (Semrush)

64%
of 47 analyzed LLMs (2019-2023) trained on filtered Common Crawl data
80%+
of GPT-3's training tokens reportedly drawn from filtered Common Crawl sources
7.8%
Wikipedia's share of ChatGPT citations, per Profound's 680M-citation analysis
6.6%
Reddit's share of Perplexity citations, same analysis

Why Wikipedia breaks the tidy version of this story

I'd love to hand you a clean, closed argument here — training data mirrors the web graph, the web graph favors a few domains, case closed, structural not preferential. The metehan.ai analysis is more honest than that, and it's worth sitting with the part that doesn't fit neatly.

In Common Crawl's own WebGraph Harmonic Centrality rankings, Facebook sits at #1, Google at #3, YouTube at #6. That tracks with what you'd expect. Wikipedia, though, ranks only #14 in HC-rank and a striking #37 in PageRank — yet it's one of ChatGPT's most-cited sources, full stop, across nearly every independent study anyone has run. If Common Crawl rank alone explained citation behavior, Wikipedia should be a mid-tier player in AI answers. It isn't.

DOMAINCOMMON CRAWL HC-RANKCOMMON CRAWL PAGERANKAI CITATION REALITY
Facebook#1Rarely a leading AI citation source
Google#3Consistently high across engines
YouTube#6Consistently high across engines
Wikipedia#14#37One of ChatGPT's most-cited sources, despite the low rank

The metehan.ai piece is careful about what this means, and I want to be equally careful repeating it: this is correlation, not full causation. Common Crawl rank is a contributing signal, not the whole mechanism. The same analysis names freshness, semantic relevance, and real-time retrieval performance as confirmed contributing factors alongside it — which is exactly why Wikipedia, a site that's constantly updated, tightly structured, and cross-referenced by nearly everything else on the web, can outperform its raw web-graph position. Structural authority gets you into consideration. Freshness, structure, and corroboration are what get you actually cited once you're there.

Wikipedia ranks #14 in HC-rank and #37 in PageRank, yet it's one of ChatGPT's most-cited sources — a case where crawl-graph rank alone doesn't fully explain citation behavior.

The question worth asking about ChatGPT citations

So where does that leave the Google statement? Basically right, and basically beside the point. Nobody flipped a switch for Reddit. But Reddit, Wikipedia, YouTube, and Google's own properties don't need a switch flipped, because they already occupy structural positions in the web graph that most training pipelines inherit by default, reinforced by exactly the kind of freshness and corroboration signals AI engines are built to reward. "No special preference" and "structurally advantaged" are not contradictions. They're two true sentences about two different mechanisms, and conflating them is how you end up either dismissing the whole conversation as paranoia or, worse, deciding the only move left is to chase a boost that was never being handed out in the first place.

We've made the second mistake's specific case before — go read why we told clients to stop chasing Reddit for AI citations if you want the tactical argument against treating any single platform as a guaranteed citation channel. That's not what this piece is about. This piece is about the layer underneath: why the same names keep winning regardless of what any one company says about its algorithm, and what that implies about where you should actually be spending effort.

1Structural position, not platform luckReddit and Wikipedia's citation dominance traces back to their long-standing position in the open web graph that training data was built from, not a favor from any single company's ranking team.
2Freshness and corroboration still do real workWikipedia's low Common Crawl rank versus its high citation share shows that being frequently updated and widely cross-referenced can outweigh raw graph position — which means those levers are available to you too.
3Denials and structural advantage can both be trueGoogle saying no preference rule exists doesn't resolve whether the underlying system produces preference-shaped outcomes. Read platform statements for what they narrowly deny, not for what they imply.

What this means for your ChatGPT citations

This is where I'd normally tell you to go build the exact thing Reddit already has, except you can't. You are not going to out-Reddit Reddit's fifteen-plus years of link density, and trying is how agencies burn budgets on strategies that were never going to compound. The useful move is different: understand which pieces of Reddit and Wikipedia's advantage are earnable by anyone, and go earn those specifically.

Some illustrative math, clearly labeled as illustrative and not drawn from either source study: if freshness and corroboration are doing real, measurable work even for domains sitting outside the top tier of Common Crawl rank, then a domain ranked, say, in the low hundreds of thousands on raw web-graph metrics isn't locked out of citations. It's just starting from a smaller structural base and needing to lean harder on the two levers that are actually within reach — publishing content that gets updated and cross-referenced rather than published once and abandoned, and getting corroborated by other sources engines already trust rather than only by your own domain.

That's the practical translation of the three signals behind most AI citations: extractable structure, demonstrated authority, and machine access. None of those three require Reddit-scale link density. They require the same boring, compounding discipline that built Reddit's position in the first place, just compressed into a much shorter runway because you're not starting in 2005. If you want the fastest earnable version of demonstrated authority, it usually runs through the specific link types AI engines actually trust rather than general backlink volume — G2 reviews, industry associations, primary research, the sources that show up as corroboration in citation data regardless of Common Crawl rank.

For B2B SaaS teams specifically, this argument matters more than it sounds like it should, because the buying committee researching your category is running exactly the kind of comparison queries where citation concentration bites hardest. We've seen this play out directly in our generative engine optimization work with clients in B2B SaaS: the teams that stopped asking "how do we get Reddit-level treatment" and started asking "what structural and freshness signals can we actually build this quarter" saw citation gains that held up across model updates, because they weren't dependent on any one platform's goodwill. Zenity's work with us on citation visibility as a B2B agentic AI security platform is one version of that: not a Reddit strategy, a structure-and-corroboration strategy.

DO THIS NEXTStop debating whether Google or any AI engine hands Reddit a secret boost. Pull your last twenty tracked ChatGPT citation gaps, and for each domain that keeps beating you, check three things: is it updated more often than your equivalent page, is it cited or linked by other sources the engine already trusts, and is your version actually structured for extraction. Fix those three before you spend another planning cycle relitigating a debate that was never going to change your citation rate either way.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

JB
Josh BernsteinMANAGING PARTNER, SOMETHING INC.

Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.