For most of the last two years, the content-strategy conversation about AI writing has been defensive: will Google penalize it, will readers trust it less, should a team disclose when a draft started as a machine output. A new Ahrefs dataset reframes the question entirely, and reframes it in a direction almost nobody was arguing for. It isn't asking whether AI content survives in search. It's showing that AI Overviews specifically cite it more than the open web contains it in the first place, which is a different and considerably more useful thing to actually understand.
The study, and how it classified content
Si Quan Ong, an SEO and marketing educator at Ahrefs, published the analysis with contributor Xibeijia Guan. The methodology: one million SERPs displaying AI Overviews, 1.9 million cited URLs extracted from the top three citations on each, and 500,000 of those URLs run through the Page Inspect feature in Ahrefs' Site Explorer, which includes a proprietary AI-content detector that scores a page on a spectrum rather than a binary label.
| CLASSIFICATION | SHARE OF AI OVERVIEW CITATIONS | SHARE OF GENERAL WEB (900K-PAGE BASELINE) |
|---|---|---|
| Pure AI | 3.6% | 2.5% |
| Pure human | 8.6% | 25.8% |
| Mixed AI and human | 87.8% | 71.7% |
The spectrum classification here matters more than the topline numbers alone, because it avoids the false binary that most AI-content coverage still leans on. "Mixed" isn't a vague middle category; Ahrefs breaks it into bands, from minimal AI use in the 1-10% range up through dominant AI use at 71-99%. The largest single band among cited pages was moderate AI use, 11-40%, at 44% of the mixed-content group. That's a page that reads as substantively human-directed with AI assistance woven through drafting or research, not a page a model wrote unsupervised and nobody touched.
The gap between citations and the open web
The comparison that makes this data worth a full piece is the baseline. Pure human content makes up 25.8% of the general web, per Ahrefs' earlier 900,000-page study, but only 8.6% of what AI Overviews actually cite. That's a threefold gap in the wrong direction for anyone who assumed unassisted human writing carries some inherent citation advantage. Mixed content, by contrast, is overrepresented in citations relative to its share of the open web: 87.8% of citations versus 71.7% of the general web.
Content classification: AI Overview citations vs. the general web baseline.
Read plainly, that's a citation system that skews toward AI-assisted pages relative to what actually exists to be cited. It doesn't mean AI Overviews are indifferent to quality, and it doesn't mean unassisted human writing is being actively filtered out. It means whatever AI-assisted content tends to do well, which the data suggests is breadth of coverage, structural completeness, and currency, happens to be the thing AI Overviews' retrieval and summarization process rewards, independent of who or what typed the words.
Why this isn't the AI-content penalty story
It's worth being precise about what this data doesn't say, because the easy misreading runs in a predictable direction: "Google prefers AI content, so write more of it, faster." That's not what 87.8% mixed-content citation share supports. We've already covered Ahrefs' 331,000-page study showing Google's ranking systems don't detect or penalize AI-written content at all, they penalize thin, unhelpful content regardless of who or what produced it. This new citation data is the companion finding, one layer further into the funnel: not just that AI content isn't punished, but that AI-assisted content is disproportionately represented in what actually gets surfaced as an authoritative source.
The correlation-versus-causation gap here is real and worth taking seriously rather than waving away. AI-assisted production and AI Overview citation could both be downstream of a third factor, teams with the resources and process discipline to produce comprehensive, well-structured, frequently updated content are also more likely to be using AI tooling somewhere in their workflow, simply because that's how competent content operations work in 2026. The AI assistance might be a marker of process quality rather than the cause of the citation, and confusing the marker for the cause is exactly the mistake this data invites if you stop reading at the headline number.
There's a second, more mundane possibility worth naming too: scale. A single writer producing careful, unassisted long-form content might publish a handful of pages a month. A team using AI tooling for research synthesis, first drafts, or update generation can plausibly cover far more of a topic cluster in the same window. If AI Overviews' retrieval process rewards a domain that has comprehensively mapped a topic space, sheer publishing velocity, not any per-page quality difference, could be doing a meaningful share of the work in this data. Nothing in Ahrefs' methodology isolates velocity from per-page quality, which means this explanation and the coverage-and-structure explanation above aren't competing theories. They're probably both true, feeding the same outcome from different angles.
What's actually driving the skew
“The AI content isn't winning citations because it's AI. It's winning because the workflow that produces it also happens to produce the coverage, freshness, and structure the citation system was already rewarding, long before any of this tooling ever existed.”
What to do with this if you write content for a living
The actionable read is not "use more AI in drafting," because that instruction skips past the actual mechanism this data points to. It's "match the properties that correlate with citation, however you produce them." A fully human-written page that covers a topic comprehensively, stays current, and uses clear, extractable structure should perform just as well as a mixed-production page with the same properties, because nothing in this dataset measures authorship directly. It measures a detector's confidence score, which is a proxy for production method, not a verified ground truth, and Ahrefs' own detector, like every AI-content classifier currently available, carries real error rates on edited and hybrid text.
This also means teams currently avoiding AI tooling out of a fear of some invisible penalty are making a decision based on a risk that this data, and the 331,000-page study before it, both fail to find any evidence for. The actual risk was never the tooling. It's publishing less comprehensively, updating less often, and structuring less clearly than a competitor who happens to be using AI assistance to do more of all three. Avoiding the tools doesn't remove that risk. It just makes the gap harder to close.
There's a broader point here about where content strategy attention should go in 2026, and it connects to the comparison-content citation advantage we've documented separately: the formats and properties that earn citations keep turning out to be structural and coverage-based, not stylistic. Whether a page was drafted with AI assistance is, per this data, a weaker predictor of citation than whether it actually answers the full scope of what someone would ask about the topic.
It's worth stating the limits of this dataset plainly too, because a study this citable will get oversimplified fast in secondhand summaries. This is one detector's classification, run against one engine's citation behavior, over one measurement window. Ahrefs doesn't claim its detector achieves perfect accuracy, and a page confidently classified as heavily AI-assisted could, in a meaningful share of cases, be a human-written page that happens to share stylistic patterns with AI output, formulaic transitions, certain sentence-length distributions, the kind of false positive every current detector produces at some rate. None of that undermines the headline gap between citation share and web-baseline share, which is large enough to survive a reasonable margin of classifier error. It does mean nobody should treat the 3.6%-versus-8.6% pure-category numbers as more precise than they actually are.
Do this next: pick your three most important pages for a topic you want to own in AI Overviews, and audit each one against the properties this data actually measures, not authorship. Does it cover the topic's full scope, including the adjacent questions a reader would ask next? Has it been meaningfully updated in the last quarter, not just date-stamped? Is the structure extractable, with a clear answer near the top of each section? Fix those three things regardless of who or what is doing the drafting, and the citation gap this study describes stops being about AI at all, and starts being about the same coverage-and-currency discipline that was always the actual job. Our content marketing team builds exactly this kind of coverage-and-cadence audit into every engagement, because the production method was never the lever that mattered.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.