This is not a study Something Inc. ran. It is something we think the industry needs more of anyway: a synthesis of three independently published 2026 datasets, read together, to answer a specific question with actual evidence instead of habit. Which GEO investments correlate with getting cited by an AI engine, and which ones are answer engine optimization tactics that agencies keep billing for out of inertia? The three sources measure different things. Read side by side, they rank cleanly into three tiers, and the ranking does not match where most enterprise budgets currently sit.
Methodology: Cross-Referencing Three 2026 Answer Engine Optimization Studies
Every number in this piece traces back to one of three sources, and none of them is ours. This is a synthesis piece, not primary research Something Inc. fielded itself, and we want that stated plainly before anything else. What we did instead: we read three independently published 2026 datasets side by side, each built on a different methodology and a different underlying dataset, because each one measures a different slice of what makes a brand show up inside an AI answer. One measures statistical correlation between off-site signals and citation outcomes. One measures where citations actually come from once they happen. One measures whether a specific technical artifact ever gets requested at all. None of the three, alone, tells the whole story. Together, they do.
The first source is Machine Relations Research's "AI Search Citation Factors" study, authored by Jaxon Parrott and published July 4, 2026. It runs Spearman correlation analysis between off-site brand signals and AI citation and visibility outcomes, using underlying data from an Ahrefs study of 75,000 brands with a Domain Rating over 40, filtered to keywords with 800 or more monthly search volume. The same research names three corroborating datasets pointing the same direction: Semrush's analysis of 126 million U.S. AI search prompts, ConvertMate's review of 80 million citations, and Digital Bloom's tracking of 680 million citations. Four independent measurement efforts, converging on the same broad pattern, is a stronger evidentiary base than any one of them alone.
The second source is Muck Rack's "What Is AI Reading?" study, now in its third edition, last updated May 7, 2026 and authored by VP Communications Linda Zebian. It analyzed more than 25 million links across ChatGPT, Claude, and Gemini responses spanning 17 industries, and it answers a different question than Parrott's correlation work: not which signals predict citation, but which categories of source actually get cited once an AI engine produces an answer. Muck Rack cofounder and CEO Greg Galant frames the finding as validation, not surprise, for communications teams already doing the work of earning coverage the hard way.
The third source is Ahrefs' llms.txt readership study, published June 15, 2026 and authored by Louise Linehan and Xibeijia Guan. It used Ahrefs' own Web Analytics and Bot Analytics products to examine server logs across 137,210 domains during May 2026, checking not whether sites publish an llms.txt file, but whether anything ever actually requests it. That question, whether a popular GEO checklist item does anything measurable once it's live, sits outside what either of the first two studies could tell you, and it's the piece that turns this synthesis from two studies into a genuine cross-check on where GEO budgets actually go.
Three studies, three different measurement methods, arriving at compatible rather than contradictory conclusions, is what synthesis is for. If only one study existed, this would be a single report and worth treating cautiously. Three, cross-referenced and pointing the same direction, is closer to a foundation we're comfortable building budget decisions against.
It's also worth naming what convergence across these three sources rules out. If Machine Relations Research's correlation ranking were an artifact of one analyst's methodology, Muck Rack's citation-share breakdown, built from an entirely different sample and a completely different counting method, wouldn't need to land anywhere near the same conclusion. It does. Earned coverage and branded presence out-perform paid and transactional tactics in both datasets, measured two different ways. That kind of agreement between independently run research teams, using different tools and different underlying data, is the strongest kind of confirmation available short of running the study ourselves.
Tier 1: The Signals With Real Correlation Strength
Three investments show up at the top of every dataset we cross-referenced, and none of them is the one most link-building retainers are scoped to deliver. YouTube mentions correlate with AI citation and visibility outcomes at 0.737, the single strongest relationship Parrott's analysis found. Branded web mentions come in at 0.664. And earned media broadly, the category Muck Rack measured directly rather than by correlation, accounts for 84% of all AI citations across the 25 million-plus links it reviewed. Three different measurements, two different studies, one consistent answer: presence and reputation outside your own domain outweighs anything you can buy by the link.
YouTube's position at the top of the correlation table is easy to miss if your GEO program is scoped purely around web content and outreach. A mention doesn't need a backlink attached to correlate with citation, and it doesn't need to originate on a channel you own. Independent creators, review channels, and comparison videos your team never touched are already producing the single strongest correlate in this entire dataset, whether or not anyone on staff is tracking it as a GEO input at all.
Spearman correlation with AI citation and visibility, by off-site signal (Machine Relations Research, Jul 2026)
That 84% figure breaks down further, and the breakdown is where the recommendation gets sharper. Journalism specifically accounts for 27% of cited sources on its own, a subset of the earned-media share rather than a separate category. Paid and advertorial content, at the other end, accounts for just 0.3% of citations, a rounding error against everything above it. If your GEO program still routes discretionary budget toward sponsored placements or advertorial content because that channel is easier to guarantee and easier to bill, this is the number that should end that conversation.
Where that earned coverage needs to land also depends on which engine your buyers actually use, a distinction our own prior citation research found holds across engines too. Muck Rack's data breaks citation behavior down by platform, and the differences are large enough to change how you prioritize outreach.
| ENGINE | CITES IN THIS SHARE OF RESPONSES | AVG. CITATIONS PER RESPONSE |
|---|---|---|
| ChatGPT | 96% | 5 |
| Gemini | 82% | 8 |
| Claude | 55% | 13 |
ChatGPT cites almost universally but shallow, five sources on average. Claude cites in barely half its responses, but when it does, it goes deep, thirteen citations on average, more than twice ChatGPT's rate. If your buyers live inside Claude conversations, getting cited at all is the harder and more valuable win; if they live in ChatGPT, breadth of coverage across more sources matters more than depth in any single one.
Tier 2: Moderate Signals Worth Secondary Budget
Below the top tier sits a set of signals that are real, measurable, and worth funding, just not worth funding first. Branded anchor text correlates at 0.527, roughly a third weaker than branded web mentions and meaningfully behind the two strongest signals in the dataset. Brand search volume sits in the 0.334-0.392 range depending on the specific query set measured. Both are genuine, both move with citation outcomes, and both are secondary to the presence signals above them, not competitors to them.
Content structure belongs in this tier too, on its own axis, separate from any off-site signal. Parrott's research found content built around 19 or more discrete data points earns two to three times more AI citations than text-only content covering the same topic, a finding our review of the broader 2026 citation data treats as one of the more actionable structural levers available, because unlike branded mentions it's entirely inside a content team's control. The same research carries a separate warning for anyone still optimizing purely for classic rankings: only 38% of AI Overview citations now come from Google's organic top 10 results, which means ranking well and getting cited well are correlated but no longer the same job, a gap our analysis of domain concentration in AI answers explores from a different angle.
For a team with a constrained GEO budget, Tier 2 is often where the marginal dollar does the most good, not because the correlation is stronger than Tier 1's, it isn't, but because it's cheaper to move. Restructuring existing content around explicit data points is largely an editorial exercise, not a new spend line, and brand search volume responds to demand generation work most marketing teams are already funding for reasons that have nothing to do with GEO. Treat Tier 2 as the low-cost adjustments made while the larger, slower Tier 1 investments in earned coverage and branded presence compound in the background.
Tier 3: Where Most Answer Engine Optimization Budgets Still Sit
At the bottom of the correlation table sits the asset most link-building retainers are literally scoped to acquire: backlinks, correlating with AI citation at just 0.218, roughly a third the strength of branded web mentions and less than a third the strength of YouTube mentions. That doesn't make backlinks worthless. Domain Rating still tracks with classic organic visibility, and link building still has a real job in classic search. Buying backlinks for SEO purposes is a different exercise than buying them for AI citation, and conflating the two is exactly the mistake this data corrects.
The second Tier 3 line item is llms.txt. Ahrefs found 28% of the 137,210 domains it scanned, 38,360 sites, publish a valid llms.txt file. Of those, 97% received zero requests during the entire month measured. Of the requests that did land, 96% came from bots generally, but named AI tools accounted for only 19.5% of that traffic, and GEO or AEO auditing tools accounted for another 12%, meaning a meaningful share of the thin traffic the file does get isn't even the audience it was built for. Ahrefs' own conclusion: "If you publish an llms.txt file today, the most likely outcome by far is that nothing ever fetches it." That lines up with what we found checking for a citation lift directly: the file is close to theater for most sites publishing it.
The pattern connecting both Tier 3 items is the same: they are the two GEO-adjacent tactics that most resemble classic SEO deliverables, a link and a technical file, and they're the two easiest to staff, price, and report on inside a standard retainer. Neither is a coincidence. Agencies built processes around producing backlinks and shipping technical artifacts over a decade of optimizing for a different system, and those processes didn't disappear when the target moved from a results page to an AI answer. They got relabeled as GEO work instead of rebuilt for it.
Notice what doesn't appear as a fourth tier: content quality, technical crawlability, or site speed. Those remain necessary conditions in every one of these three studies' own framing, a page that can't be rendered or parsed doesn't get evaluated for anything above, but none of the three datasets measured them as a variable with its own correlation coefficient. That's a genuine gap in what currently exists to measure, not evidence that those fundamentals don't matter. Read the three tiers above as a ranking of investments beyond the technical floor, not a replacement for it.
What This Synthesis Can't Tell You
Read the tiers above as strength of relationship, not as a guarantee of causation. Parrott's Spearman correlations describe how strongly two variables move together across 75,000 brands, not that funding YouTube mentions directly causes a citation the following week. Muck Rack's 84% figure describes where citations already land, not what happens if a brand with zero earned coverage starts investing in it today. Ahrefs' llms.txt numbers describe one May 2026 snapshot of server logs, not a permanent verdict on the format going forward. None of that undermines the ranking; a correlation of 0.737 held across 75,000 brands and four corroborating datasets is not noise, and 97% of files getting zero requests is not a fluke. But treat this framework as the best available evidence ranking, not as a promise that Tier 1 spend converts to citations on a fixed timeline.
There's a sampling caveat worth naming plainly, too. Machine Relations Research's underlying Ahrefs data filtered to brands with Domain Rating over 40 and keywords with 800 or more monthly searches, which skews the sample toward already-established sites with existing search visibility. Muck Rack's 25 million-plus link sample spans 17 industries but is necessarily built from links AI engines already chose to cite, not a random sample of the whole web. Neither limitation invalidates the ranking; it means the ranking describes what predicts citation among sites already competing for it, not what it takes for a brand starting from zero to break in for the first time.
The Verdict: Budgets Are Inherited From SEO, Not From This Data
Line the three tiers up against where enterprise GEO budgets actually sit today, and the mismatch is the whole finding. Backlink acquisition and technical checkbox items like llms.txt still absorb a disproportionate share of GEO spend at most organizations we work with, and the reason isn't that the 2026 data supports that weighting. It's that those are the tactics inherited wholesale from classic SEO programs, where a link was the deliverable and a technical audit was the quarterly report. Nothing about that inheritance was updated when the goal shifted from ranking on a results page to getting named inside an answer. Generative engine optimization done to the 2026 evidence looks different: budget weighted toward earned coverage and branded presence first, structured and data-dense content second, and backlinks and llms.txt maintained as low-cost, low-priority line items rather than headline deliverables. This matters most for categories like B2B SaaS, where the buying committee already treats third-party validation as part of its own research process before a vendor's own site enters the picture at all, and where measuring citation and mention market share directly is the only way to know whether a rebalanced budget is actually working.
None of this is an argument for abandoning technical GEO work or ending link-building programs outright. It's an argument for sequencing: the technical floor first, Tier 1 investment next, Tier 2 refinements alongside it, and Tier 3 tactics maintained at the minimum viable level rather than treated as the headline of the retainer.
Audit Your GEO Budget Against This Framework
For illustration only: a program spending $10,000 a month with 60% routed to backlink acquisition and llms.txt maintenance, and 40% to earned media and branded presence, is running roughly the inverse of what Tier 1 and Tier 3 evidence would suggest. Flipping that ratio doesn't require a bigger budget, only a different allocation of the one already approved.
Pull your GEO and link-building line items for the current quarter and sort every dollar into one of the three tiers above. If Tier 3 (backlink acquisition, llms.txt maintenance, technical checkbox work) accounts for more than a third of the total, that's the specific number to bring into your next budget review, not a general sense that things could be better. Redirect the difference toward earned media outreach, YouTube presence, and branded mention tracking, the three inputs every dataset in this study agrees matter most, and re-run this same audit next quarter against the same three tiers to see whether the rebalancing actually moved your citation numbers.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.