Something Inc.LoginSchedule a free consultation
LINK BUILDING

Link building for AI citations: the 84% earned media number

Muck Rack analysed 25 million cited links and found 84% traced to earned media. A separate dataset shows documentation surging in ChatGPT. Both are true, and the reason why decides your budget.

JBJosh BernsteinManaging Partner · AUG 24, 2026 · 12 MIN READ

Muck Rack's Generative Pulse team published the third edition of its citation analysis in May, built on more than 25 million cited links across 17 industries and three engines. The headline number is that 84% of AI citations trace to earned media rather than owned or paid content. That figure has held between 82% and 89% across all three editions, which is the part that makes it worth planning against rather than reacting to. Link building for AI citations is a different discipline from link building for rankings, and this is the dataset that shows how different.

TL;DR · 60 SECONDSEarned media accounts for 84% of AI citations, journalism specifically for 27%, and paid or advertorial content for 0.3%. Meanwhile a separate August dataset shows help centers and documentation surging in ChatGPT's cited sources. The two findings look contradictory and are not: they describe different query types. Getting that distinction right is what turns this from a headline into a budget.

Take the composition first. Of everything cited across the sample, 84% is earned media: coverage, mentions and references a brand did not publish and did not pay for. Journalism proper accounts for 27% of cited sources, spread across more than 20,000 distinct outlets, which is a wider tail than most digital PR strategies assume when they build a target list of forty publications. Paid and advertorial content accounts for 0.3%. That last number deserves a moment on its own.

Three tenths of one percent. Whatever sponsored content is doing for you, and it may be doing real things for brand and for direct traffic, it is not appearing in the pool of documents AI engines cite. If your visibility plan for generative surfaces runs through paid placements, this dataset says that plan is aimed at a channel representing a rounding error of citations. That is a clean, actionable finding and it costs nothing to act on: stop counting sponsored placements as AI visibility work and count them as whatever they actually are.

The consistency across editions is what makes the composition credible. Earned media has stayed in an 82% to 89% band and journalism in a 25% to 27% band across three separate measurement periods running back to July 2025. Single studies in this space contradict each other constantly, usually because prompt sets differ. A number that holds its shape across three collection windows is measuring something structural. Muck Rack publishes the round-up openly as what AI is reading, and the methodology note is worth the click: the sample covers ChatGPT, Claude and Gemini, so nothing here describes Perplexity or Google's AI Overviews, and treating it as a claim about all AI search would be overreading it.

The per-engine breakdown is where the strategy actually gets set, because the three engines behave so differently that averaging them produces advice that fits none of them.

ENGINESHARE OF RESPONSES WITH CITATIONSAVERAGE CITATIONS PER RESPONSEMOST CITED DOMAIN
ChatGPT96%5Wikipedia
Gemini82%8Reddit
Claude55%13PubMed Central

Read that table as an inverse relationship and it starts to make sense. ChatGPT cites on almost every response but cites few sources each time, which makes each slot extremely competitive. Claude cites on barely half its responses but when it does, it cites thirteen sources, which makes any individual slot much easier to win and much less valuable. Gemini sits between them. A brand that is cited once per response in ChatGPT is in a materially stronger position than a brand cited once per response in Claude, and any dashboard that reports a single blended citation count across engines is destroying that distinction before you see it.

The top-domain column tells you what each engine reaches for by default. Wikipedia for ChatGPT, Reddit for Gemini, PubMed Central for Claude. Those are three completely different theories of authority: encyclopedic consensus, community discussion, peer-reviewed literature. Note also that Gemini's reliance on Reddit is the counterpoint to what happened inside ChatGPT this month, where community threads collapsed as a citation source while Google's surfaces barely moved. Engine-level divergence is not noise. It is the main feature of the current environment. It also means the phrase getting cited by AI is close to meaningless without a named engine attached to it. Winning a slot in Claude, where thirteen sources are listed and only half of responses cite anything at all, is a fundamentally different achievement from winning one of five slots in ChatGPT, which cites on 96% of responses. Those are not the same goal, they do not respond to the same work, and a client brief that asks for both without distinguishing them is asking for two projects.

One more finding from the same dataset is worth carrying around: Axios appears in ChatGPT's top three cited domains in 13 of 17 industries. Not in its own industry. In thirteen of them. That is a single outlet functioning as a general-purpose authority across most of the economy, which tells you something uncomfortable and useful about how much of the citation pool is decided by a handful of publications with broad topical coverage and clean, scannable formatting.

The correlation data that reorders your budget

Composition tells you what gets cited. It does not tell you what predicts whether your brand gets cited. For that, the useful reference is Ahrefs' correlation work across roughly 75,000 brands with a Domain Rating above 40, measured on keywords with at least 800 monthly searches, using Spearman correlations against AI visibility.

YouTube mentions (0.737)74%
Branded web mentions (0.660 to 0.710)69%
Branded anchor text (0.511 to 0.628)57%
Brand search volume (0.352 to 0.466)41%
Domain Rating (0.266 to 0.326)30%
Backlink count (0.218)22%

Spearman correlation with AI visibility, per Ahrefs' analysis of about 75,000 brands (correlation, not causation)

Backlink count sits at the bottom at 0.218, roughly a third of the correlation of branded web mentions. That is the finding that should change how a link building budget is written, and it needs stating carefully, because correlation is not causation and none of these figures prove that acquiring mentions causes citations. Large brands have more of everything. What the ranking does establish is that a program optimizing purely for link volume is optimizing the weakest signal in the set, and one that produces mentions as a side effect of coverage is picking up the stronger ones for free.

The complementary figure is that only about 38% of AI Overview citations come from Google's own top ten organic results. If you have been assuming that ranking well is a reliable path into AI citations, that number sets the ceiling on the assumption. Roughly six in ten citations are going to pages that did not win the classic ranking contest for that query. The YouTube correlation is the one we have written about separately in YouTube mentions as a citation signal, and it remains the single strongest number in the set.

Reconciling earned media with the documentation surge

Now the apparent contradiction, and the reason this piece exists. Earlier this month, a separate source-category comparison showed help centers and product documentation moving from roughly 2% of ChatGPT's cited sources to 32%, with smaller company sites and blogs halving. That is first-party, owned content surging. Muck Rack says 84% of citations are earned. Both cannot be describing the same thing.

They are not. They are describing different questions, and Muck Rack's own data contains the reconciliation. In their sample, industry trend queries cite journalism at roughly double the rate of how-to questions, and press releases appear about 3.5 times more often in trend responses than in best-of queries. The engine is not applying one theory of authority. It is choosing source types according to what the question needs.

Which is obvious once stated. Ask what is happening in a market and the appropriate source is journalism, because a vendor cannot credibly narrate its own category. Ask how to configure a product and the appropriate source is that product's documentation, because nobody else knows. Ask which tool is best and the engine reaches for review platforms, comparison content and community discussion. The 84% and the 32% are answers to different questions, and the mistake is reading either as a universal ranking of source types.

EARNED
Trend and market questionsWhat is happening in a category, what changed, who is growing. Journalism dominates, press releases carry 3.5x their usual weight, and earned media is close to the only route in.
OWNED
How-to and configuration questionsHow to set something up, integrate it, or fix it. Documentation and help centers own this, and no amount of press coverage substitutes for a clear, crawlable docs page.
MIXED
Comparison and shortlist questionsWhich tool is best for a use case. Review platforms, comparison pages and community threads compete here, and engine behaviour on this class is the least stable of the three.

Map your own demand onto those three buckets before allocating anything. Most B2B software companies find their commercially valuable questions cluster heavily in buckets two and three, with bucket one mattering for category creation and for the analyst-adjacent audience rather than for pipeline. A company in a genuinely new category has the opposite distribution. The allocation follows the distribution, and it is different per client, which is why a single answer to what earns AI citations is always wrong. The failure mode we see most often is a company reading one strong study, recognising it as credible, and rebuilding its entire program around a distribution that belongs to somebody else's category. Both datasets discussed here are good. Neither was measured on your questions.

Here is the operating version. Run the twenty to fifty questions that matter commercially through the engines you care about, and record the source type of every citation returned. Not the domain. The type: journalism, documentation, review platform, community, marketplace, encyclopedic. That classification takes an afternoon and it produces the only allocation input that reflects your actual category rather than an industry average.

Then spend against the distribution you measured. If journalism is carrying half your citations, digital PR is your AI visibility program and should be funded as such rather than as a link acquisition line item. If documentation is carrying half, your budget belongs with whoever owns the help center, which is usually not marketing at all. If review platforms carry half, the work is profile completeness and review volume. Most programs will land on a split, and the split is the plan. We laid out the durability side of this trade in the link building budget decision table, and it pairs directly with the allocation logic here.

Two things to stop doing while you are at it. Stop counting sponsored placements as AI visibility, given they represent 0.3% of citations. And stop reporting a single blended citation number across engines that cite at 96%, 82% and 55% rates with five, eight and thirteen sources per response. Split it by engine or the number means nothing. That per-engine discipline is how we run link building and digital PR engagements for cybersecurity and technical clients, where the gap between what one engine rewards and what another does is wide enough to change the whole plan.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

JB
Josh BernsteinMANAGING PARTNER, SOMETHING INC.

Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.