Something Inc.Schedule a free consultation
RESEARCH

What actually moves AI citations: the 2026 evidence review

We graded eight widely sold GEO interventions by the strength of the evidence behind them. The two with real experimental designs both returned null results, and everything with a positive finding is correlational.

8 INTERVENTIONSEVIDENCE GRADEDAUG 2026
8
interventions assessed against published 2026 evidence
2
with a genuine experimental or quasi-experimental design
0
of those two that produced a positive effect
6
supported only by correlation or description
ABSTRACTGenerative engine optimization is sold as a set of interventions, but the AI citation evidence behind those interventions varies enormously in quality and almost nobody grades it. We assessed eight commonly recommended tactics against the published 2026 research, sorting each by study design rather than by headline. Two interventions have been tested with designs capable of supporting a causal claim, one randomized A/B test and one difference-in-differences quasi-experiment. Both returned null results. The six interventions with positive findings are supported by correlational or descriptive work only, several of it strong and large-sample, none of it capable of establishing that doing the thing causes the citation. The practical conclusion is not that GEO does not work. It is that the industry's confidence is distributed almost exactly inversely to its evidence.

Every agency deck in this category contains a list of things to do. Publish llms.txt. Add schema. Serve markdown to crawlers. Earn brand mentions. Refresh old pages. Front-load the answer. The lists are remarkably consistent across firms, which looks like consensus and is closer to circulation. Very few of those recommendations have been tested, and the two that have been tested properly did not survive it.

What follows sorts the published AI citation evidence by what each study was actually capable of showing. Nothing here is new data collection. It is a reading of other people's work, with the design of each study treated as the primary fact about it.

Methodology

We selected interventions that appear repeatedly in commercial GEO recommendations and for which at least one public study exists with a stated sample and method. For each, we recorded the strongest available study, its design, sample size, collection window, effect direction and magnitude, and the caveats stated by its own authors. Studies without a disclosed sample size or method were excluded regardless of how widely they are quoted, which removed several frequently cited claims about earned media share of citations that trace back to no verifiable primary source.

Interventions were then graded into three tiers by design strength, not by result. A null finding from a strong design ranks above a large positive finding from a weak one. That ordering is standard in every field that takes evidence seriously and almost unheard of in this one.

TIERDESIGNWHAT IT CAN SUPPORT
Tier 1Randomized A/B or difference-in-differences with controlsA causal claim, within the tested population
Tier 2Large-sample correlation across many brands or pagesAssociation, and a prioritization argument
Tier 3Descriptive analysis of observed citationsPattern description, and hypothesis generation
ExcludedUndisclosed method or sample, single-client anecdoteNothing

The evidence hierarchy

The distribution is the finding. Of the eight interventions assessed, two sit in tier one, three in tier two, and three in tier three. Both tier one results are null. Every positive result in the set comes from tier two or tier three.

INTERVENTIONSTRONGEST STUDYDESIGNRESULT
Serve markdown to AI crawlersProfound, Feb 26, 2026Tier 1, randomized A/BNull
Add JSON-LD schemaAhrefs, May 11, 2026Tier 1, difference-in-differencesNull to slightly negative
Publish llms.txtSE Ranking and Trakkr, 2026Tier 2, large-sample observationNo measurable lift
Earn brand and video mentionsAhrefs, Dec 12, 2025Tier 2, correlation, n=75,000Strong association
Rank in classic searchAirOps and Indig, Mar 2026Tier 2, correlationStrong association
Refresh existing pagesSeer Interactive, Jul 24, 2026Tier 2, correlationModerate association
Front-load the answerIndig, Feb 18, 2026Tier 3, descriptiveContradicts the advice
Write longer pagesIndig, Mar 24, 2026Tier 3, descriptivePositive association

Tier 1: controlled tests, and what they found

Only two published tests in this category have a design capable of isolating an effect. Both deserve to be read in full by anyone selling either tactic.

Profound ran a randomized A/B test of serving markdown versions of pages to AI crawlers, across 381 pages on six websites over three weeks. The result was a mean lift of roughly 16% that was not statistically significant, with a median impact of about one additional visit over the entire three-week window. The experiment was powered to detect effects above 40% and found nothing at that magnitude. The same test recorded the bot mix hitting those pages: ChatGPT-User at 73%, Meta at 20%, OAI-SearchBot at 4%, ClaudeBot at 2%, and GPTBot at 1%. That distribution is arguably more useful than the headline, because it says the on-demand retrieval agent, not the training crawler, is what most sites are actually serving.

Ahrefs tested schema with a difference-in-differences design: 1,885 pages that added JSON-LD, matched against 4,000 control pages, over August 2025 to March 2026. The measured effect on AI citations was minus 4.6% on Google AI Overviews, a small but statistically significant decline, plus 2.4% on Google AI Mode and plus 2.2% on ChatGPT, both indistinguishable from zero. The authors state the limiting caveat themselves: the pages studied already carried more than 100 citations before treatment, so this measures uplift on already-cited pages rather than discovery for uncited ones.

1Null is not nothingA well-powered null result is informative. It bounds the effect. Profound's test says markdown mirrors do not produce a 40% lift, which is exactly the size of claim being made for them.
2Both tests measured uplift, not entryNeither design tells you whether these tactics help a page that is currently invisible. That is a real gap and the honest reading of both papers acknowledges it.
3Nobody has replicated either oneTwo studies, two teams, no replication. In any other field that would be treated as preliminary. Here it is the strongest evidence available.

The schema result deserves particular attention because schema is among the most commonly sold GEO deliverables. It also converges with a separate mechanical argument we covered when schema markup was found not to be the citation lever the industry markets: structured data is frequently stripped during text extraction, so there is a plausible reason the measured effect is zero. When a null result has a mechanism behind it, it should update your beliefs more than a null result that arrives unexplained.

Tier 2: large-sample correlations

Correlation cannot establish causation, and in this category the confounding is severe: brands with more of everything tend to have more of everything. But large-sample correlation is still the best available guide to where to look, and one of these datasets is large enough to be worth taking seriously.

Ahrefs computed Spearman correlations between various brand signals and AI visibility across 75,000 brands, filtered to domains above Domain Rating 40, in December 2025. YouTube mentions correlated at 0.737 with ChatGPT visibility, 0.712 with AI Mode, and 0.740 with AI Overviews. Branded web mentions came in at 0.664, 0.709, and 0.656. Branded anchors ranged from 0.511 to 0.628. Branded search volume sat between 0.352 and 0.466. Domain Rating managed 0.266 to 0.326. Raw backlink counts and URL Rating showed, in the authors' words, very weak correlations across all three surfaces.

SIGNALCHATGPTAI MODEAI OVERVIEWS
YouTube mentions0.7370.7120.740
Branded web mentions0.6640.7090.656
Branded anchors0.5110.6280.527
Branded search volume0.3520.4660.392
Domain Rating0.2660.2850.326
Backlink count, URL RatingVery weakVery weakVery weak

Two caveats that get dropped in secondary coverage. The data is from December 2025, and no 2026 replication at comparable scale exists. And the Domain Rating above 40 filter excludes precisely the mid-market sites that make up most agency client rosters, so the coefficients describe a population your client may not belong to. We worked through the practical implications when earned media and backlinks were compared as AI citation drivers, and the ratio claims circulating about that dataset do not appear in the original.

Classic ranking shows a similarly strong association from two independent sources. AirOps found 55.8% of cited pages ranked in Google's top 20 and that position one pages were cited about 3.5 times more often than pages outside the top 20. Kevin Indig's separate analysis of roughly 98,000 citation rows found pages ranking first were cited 43.2% of the time, again around 3.5 times the rate beyond position 20. Two datasets, two methods, the same multiple. That is the closest thing to replication in this entire set.

Cited pages ranking in Google top 20 (AirOps)56%
Position-one pages that were cited (Indig)43%
Retrieved pages that ever get cited (AirOps)15%

Share of AI-cited pages that also rank in Google's top 20, and citation rate for position-one pages.

Freshness sits in the same tier with a moderate association. Seer Interactive analysed 7,683 pages and 47,097 citations between March and June 2026 and found 75% of cited pages had been updated within a year, but only 42% were fresh by original publish date against 72% by update date. That 30-point gap is old pages that were maintained, which is why we read it as an argument for a refresh programme rather than more volume. It remains correlational, and only two thirds of pages could be reliably dated.

Tier 3: observational description

Descriptive work cannot tell you what to do, but it can tell you that a widely repeated instruction is wrong, which is a useful thing for a tier to be able to do.

Indig's February analysis matched 18,012 verified citations against 1.2 million AI answers using sentence-transformer embeddings. Position within the page mattered: 44.2% of citations came from the first 30% of the content, 31.1% from the middle, 24.7% from the final third. So far this supports front-loading. But within paragraphs the pattern inverts. Fifty three percent of cited passages came from the middle of a paragraph, against 24.5% from the first sentence and 22.5% from the last.

That directly contradicts the standard instruction to open every paragraph with the answer. What the cited passages had in common was different: they were nearly twice as likely to contain a clear definition, twice as likely to include a question mark, and averaged 20.6% proper nouns against a typical 5% to 8%. Reading grade level for the best-performing content averaged 16, against 19.1 for lower performers, which cuts against the advice to write simpler for machines and also against writing impenetrably.

DESCRIPTIVE
Density of named entities20.6% proper nouns in heavily cited passages. Name the tools, companies, standards, and people. Vague prose does not get quoted.
DESCRIPTIVE
Definitions get liftedCited passages were about twice as likely to contain a clear definition. A sentence that defines something is a sentence that can stand alone.
CONTRADICTS
Mid-paragraph, not first sentenceThe answer-first instruction is not what the citation data shows. Complete explanation beats positional formatting.
DESCRIPTIVE
Length correlates upwardPages above 20,000 characters averaged 10.18 citations against 2.39 for pages under 500. More passages, more chances.

The length finding needs care. It is almost certainly confounded, because long pages are usually the pages a team invested in, and investment brings other advantages. Read it as permission to cover a topic completely in one place rather than as an instruction to pad, and pair it with the retrieval-side finding that most citation opportunity sits behind machine-written fan-out queries with no search volume attached.

What the AI citation evidence supports

Stated as plainly as the evidence allows, and separating what we know from what we suspect.

Two things have been tested and did not work at the magnitudes claimed: markdown mirrors for crawlers, and adding schema to pages that are already cited. Neither result rules out a small effect and neither addresses discovery for uncited pages. But if you are paying for either as a citation lever, the AI citation evidence does not currently support the invoice.

One thing has been examined at scale twice and found nothing: llms.txt. SE Ranking across roughly 300,000 domains and Trakkr across 37,894 domains both reported no measurable citation lift, which is why we concluded that machine access matters and the file itself does not. Publish it if you want a cheap directory of your best pages. Do not sell it as a citation intervention.

Three things correlate strongly and are worth funding on that basis while acknowledging what a correlation is: brand and video mentions, classic search ranking, and maintaining existing pages. The ranking association is the best-supported claim in the set because two independent teams found the same 3.5x multiple.

The strongest evidence in generative engine optimization currently points at doing search engine optimization well and being a brand people talk about. That is an unglamorous conclusion, and it is the one the data supports.

What we would still fund, and why

Evidence grading is not a reason to do nothing. It is a reason to be honest about which line items are bets. Here is how we would split a generative engine optimization budget given what is currently known.

HOW A QUESTION BECOMES A CITATION
Machine accesscheap, necessary, unproven as a lever
Ranking and coveragebest-evidenced, two independent datasets
Brand and third-party mentionsstrongest correlation in the set
Maintenance over volumemoderate evidence, low cost
Measurementthe only way to test your own case

Machine access stays funded despite the null results, because the cost is trivial and because a page that crawlers cannot fetch or parse scores a guaranteed zero rather than an uncertain one. The distinction is between a necessary condition and a lever, and the industry has spent two years selling the former as if it were the latter, which is how a hygiene task ended up with a premium attached to it. Everything else follows the evidence: fund ranking and brand presence first, maintenance second, and treat structural formatting tactics as cheap experiments you run and measure rather than deliverables you charge for.

The final line item is the one most programmes skip. None of this evidence is about your site, and site-level effects in this category appear to vary enormously by vertical. Reconciling four independent AI visibility datasets produced agreement on very little except that measurement definitions drive most published disagreements. Run your own before-and-after on one change at a time, and you will know more about your own case than any of these papers can tell you.

Limitations

This is a review of other people's work and inherits every weakness in it. The two tier one studies are unreplicated, and both were run by companies selling adjacent products, which is true of almost every dataset in this category and does not by itself invalidate them. Several studies are six to nine months old in a category that changes quarterly. The Ahrefs correlation data predates 2026 entirely. Sample populations differ enough that the studies are not strictly comparable, and our tiering is a judgment about design quality rather than a formal appraisal instrument.

We also excluded a set of widely circulated claims, including several specific figures about earned media's share of AI citations, because we could not trace them to a primary study with a disclosed method. Their absence here is not evidence they are wrong. It is evidence that nobody has shown they are right.

If you want the underlying model of how a citation gets earned rather than the evidence audit, we set it out in the anatomy of an AI citation.

Cite this research

Cite this research● LIVE
Something Inc. (2026). What actually moves AI citations:
the 2026 evidence review.
Eight generative engine optimization interventions graded by
study design across published 2026 research.
somethinginc.com/blog/ai-citations-evidence-review-2026
 
Primary sources assessed:
Profound, markdown vs HTML randomized test, Feb 26, 2026
Ahrefs, schema difference-in-differences, May 11, 2026
Ahrefs, brand signal correlations, n=75,000, Dec 12, 2025
AirOps, retrieval vs citation, n=548,534 pages, Mar 13, 2026
Kevin Indig, citation passage analysis, Feb 18 and Mar 24, 2026
Seer Interactive, content recency, n=7,683 pages, Jul 24, 2026
SE Ranking and Trakkr, llms.txt citation lift, 2026

External sources: Profound's markdown test (Profound) and Ahrefs' schema experiment (Ahrefs). If a vendor quotes a citation lift number at you this quarter, ask two questions before anything else: what was the design, and what was the sample. Most of the time you will not get an answer, and that is an answer.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

TT
Tyler TruffiMANAGING PARTNER, SOMETHING INC.

Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.