Something Inc.Schedule a free consultation
RESEARCH

Which content formats AI engines cite most

We classified 6,200 URLs cited across ChatGPT, Perplexity, Claude, and Google AI Mode by content format and domain type. One format earns almost a third of every citation, and the gap between engines is wider than most teams plan for.

6,200 CITATIONS4 ENGINESQ2 2026
31.8%
of citations came from comparison and alternatives content
22.4%
came from how-to and guide content, the second largest share
0.41
median cross-engine agreement on which URL to cite (Jaccard)
2.6x
citation lift after a page was rebuilt for extraction
METHODOLOGYBetween April 1 and June 20, 2026 we ran 5,000 high-intent B2B prompts three times each across ChatGPT, Perplexity, Claude, and Google AI Mode, then collected every URL those engines cited. After removing duplicates and homepages, 6,200 unique cited URLs remained. Two reviewers independently classified each URL by content format (comparison, how-to, product and docs, research and data, community) and by domain type (owned brand, vendor-neutral publisher, community platform, reference, review platform). Inter-rater agreement was 0.88. Format shares are computed over all 6,200 URLs unless a figure states otherwise. This is observational data, not a controlled experiment, so we report shares and associations rather than causal effects.

Marketing teams keep asking a version of the same question: if an AI engine is going to summarize our category and cite a few sources, what kind of page do we need to be? Not what topic, what format. We had the citation data to answer it directly, so we classified every cited URL we could collect over a full quarter and looked at which shapes of content the engines actually reach for.

The short version is that format matters more than most content calendars assume, and it matters in a way that is stable enough to plan around. Across 6,200 cited URLs, five formats accounted for essentially all of the volume, and their ranking barely moved from month to month. But the moment you split the same data by engine, the neat picture fractures. A format that wins on one engine can be a runner-up on another, which is why a single content bet aimed at all four engines tends to underperform on at least two of them.

Share of citations by content format

Comparison and alternatives content took the largest share by a wide margin, at 31.8% of all cited URLs. This is the head-to-head page, the best-of list, and the alternatives-to-X page. It wins because it maps onto the buyer's first real question, which is almost always some form of what are my options, and because a good comparison states a verdict an engine can lift in one sentence. How-to and guide content followed at 22.4%, product and documentation at 17.6%, research and data at 15.9%, and community threads at 12.3%.

The two leading formats are not interchangeable. Comparison content wins the top of the funnel, the moment a buyer is still assembling a shortlist. How-to and guide content wins later, once someone has chosen a direction and is asking how to actually do the thing. If your library is heavy on one and thin on the other, you are visible for one half of the journey and absent for the other. The long tail, community threads at the bottom, still earned one citation in eight, which is more than most brands would guess and a reminder that engines weight corroboration from places they already trust.

What surprised us was how stable this ranking held. We recomputed the format shares for each of the three months separately, and the top-line order never changed. Comparison led every month, community trailed every month, and the largest month-over-month swing for any single format was 2.1 points. That stability is the reason we treat format as something you can plan a year of content around rather than a moving target you have to chase. The variance that does exist is concentrated across engines, not across time, which is the subject of the next section.

Comparison / alternatives31.8%
How-to and guides22.4%
Product and docs17.6%
Research and data15.9%
Community threads12.3%

Figure 1. Share of citations by content format, n=6,200 cited URLs.

The table below adds two things the bar chart hides: the kind of prompt each format tends to win, and its median citation rank, meaning where it sits in the source list when it does appear. Comparison content does not just appear most often, it appears higher in the list, which matters because engines lean hardest on the first sources they name.

CONTENT FORMATSHAREWINS ON PROMPTS LIKEMEDIAN RANK
Comparison / alternatives31.8%best X for, X vs Y, alternatives to X1.9
How-to and guides22.4%how to do X, X setup, X best practices2.4
Product and docs17.6%does X support Y, X pricing, X limits2.7
Research and data15.9%X benchmark, X statistics, state of X2.5
Community threads12.3%is X worth it, X honest review, X problems3.3

The same formats behave differently by engine

This is the finding that changes how you should spend. The overall ranking is a blended average, and no single engine looks exactly like it. Comparison content ranged from a high of 37.1% of Perplexity citations down to 26.9% of Claude citations. Perplexity behaves the most like a shortlist builder, favoring vendor-neutral comparisons and reviews. Claude spread its citations more evenly across formats and leaned harder on documentation and primary research than the other three. ChatGPT and Google AI Mode sat between them, closer to the blended average.

The practical consequence is that cross-engine agreement is low. When we asked how often two engines cited the same URL for the same prompt, the median overlap was a Jaccard index of 0.41, meaning roughly two of every five cited sources were shared and three were not. Visibility on one engine does not transfer to the next. A page tuned only for how ChatGPT summarizes will quietly miss on Perplexity, where a sourced comparison would have won instead.

FORMATCHATGPTPERPLEXITYCLAUDEAI MODE
Comparison / alternatives32.6%37.1%26.9%30.5%
How-to and guides23.1%18.4%22.8%25.2%
Product and docs16.9%15.2%21.4%16.8%
Research and data15.0%16.1%18.7%14.0%
Community threads12.4%13.2%10.2%13.5%

Read the columns, not the rows, and a strategy falls out. If your category conversations happen on Perplexity, comparison content is close to the whole game. If they happen on Claude, a well-sourced guide or a piece of original research has a much better shot than the blended average would suggest, and thin documentation is a liability rather than a filler asset. The same content budget, aimed engine by engine instead of at an imaginary average engine, buys measurably more coverage.

What makes a page extractable

Format is the shape of the answer, but within any format some pages get cited and near-identical ones do not. To isolate what separates them we took 300 pages that were competing for the same set of prompts, split them into cited and never-cited groups, and compared their structure. Four page-level attributes tracked with citation more tightly than anything else, and none of them is about topic or word count.

1A verdict in the first 120 wordsCited pages state their answer near the top, before any setup. Engines lift a self-contained claim; they rarely reward a conclusion buried under 800 words of throat-clearing. Pages that led with a direct verdict were cited 2.6x more often than matched pages that led with narrative.
2One idea per section, under a literal headingHeadings phrased as the question a buyer would ask, followed by a single answerable idea, extract cleanly. Cited pages averaged 41 words per paragraph; never-cited pages averaged 68. Shorter, self-contained blocks give the engine a clean unit to quote.
3A table or list carrying the comparisonWhen the core claim lived in a structured table or list rather than prose, citation rank improved by roughly a full position. Structure is not decoration; it is the format an engine can copy without paraphrasing and getting it wrong.
4Sourced, specific claimsCited pages backed assertions with a number, a date, or a named source. Vague, hedged claims were the single most common trait of the never-cited group. Engines prefer to quote something they can attribute cleanly, because an unsourced claim is a risk they push down the list.

The lift figure is worth sitting with. Rebuilding a page to lead with a verdict, tighten its sections, and carry its core claim in a table raised citation frequency 2.6x on the matched set. That is a structural edit, not a new page and not a link campaign. Most enterprise libraries are full of pages that already rank and already have the right facts, buried in a shape no engine can lift.

It is worth being precise about what this figure is and is not. The 2.6x is a before-and-after comparison on pages we rebuilt, holding topic and target prompt constant, so it controls for the two variables teams most often confuse with format. It is not a randomized trial, and other things moved during the observation window, so we would not defend it to the second decimal. But the direction and the rough magnitude were consistent across the pages we reworked, and the pattern lines up with what the four extractability signals predict. The honest read is that structure is a large, cheap lever that most content teams have not pulled, not that a single edit guarantees a specific multiple.

Engines do not reward the best-written page. They reward the one they can quote correctly without reading the whole thing.

Freshness moves some formats and not others

We expected freshness to matter everywhere. It does not. We grouped cited URLs by how recently they had been published or substantively updated and looked at the share of citations going to pages touched in the prior 90 days. The effect was strong for two formats and close to flat for two others, which means a blanket refresh-everything policy wastes effort on content that does not decay.

Research and data61%
Product and docs54%
Comparison / alternatives47%
How-to and guides33%
Community threads29%

Figure 2. Share of each format's citations going to pages updated in the last 90 days.

Research and data pages showed the strongest recency pull, with 61% of their citations going to material updated in the last quarter. That fits: a statistic engines cite is only useful if it is current, and a 2024 benchmark gets displaced the moment a 2026 one exists. Product and documentation followed at 54%, because pricing, limits, and capabilities change and engines have learned not to trust stale specs. Comparison content sat in the middle at 47%. How-to and community content barely moved, because the underlying answer is durable; a guide to a stable process does not become wrong because a calendar quarter passed.

One caution reading Figure 2: a high recency share does not by itself prove the engines reward freshness, because newer pages also tend to be better structured and more thoroughly sourced. To separate the two we compared same-age pages within each format, and the recency association survived for research and documentation but shrank toward noise for guides and community content. That is consistent with a simple mental model. Engines discount a claim when its truth has a shelf life, and they mostly ignore the publish date when the underlying answer does not expire.

The operational read is to match your update cadence to the format. Put your refresh budget behind research and documentation, where recency is doing real work, and stop re-dating evergreen guides in the hope it helps. It mostly does not, and the effort is better spent building the comparison pages you are missing.

Domain type sits underneath format

Format tells you what shape of page to build. Domain type tells you where an engine expects that shape to live, and the two interact. Vendor-neutral publishers and review platforms punched well above owned brand domains for comparison content, because engines treat an independent head-to-head as more credible than a vendor grading its own homework. Owned domains won cleanly on documentation and on original research, where being the primary source is the whole point.

DOMAIN TYPESHARE OF CITATIONSSTRONGEST FORMAT
Owned brand domain34.5%Product and docs
Vendor-neutral publisher26.2%Comparison / alternatives
Community platform16.8%Community threads
Reference (encyclopedic)12.1%Research and data
Review platform10.4%Comparison / alternatives

Owned domains earned the single largest slice at 34.5%, which is reassuring, but the number hides a trap. That share concentrates in documentation and self-published research, the two places where being the source is an advantage. For comparison content, the format that wins the most citations overall, engines lean toward vendor-neutral publishers and review platforms, which together took better than a third of the total. You cannot buy your way onto those with owned content alone. Winning the comparison layer means earning a credible presence on the independent sources your buyers and the engines both already trust.

What this means for you

The data points to a sequence, not a single tactic. First, decide which format your category actually needs before you brief a single piece. If your buyers ask AI engines to compare options, comparison and alternatives content is where a third of the visible citations are, and you almost certainly under-produce it. If they ask how to operate a tool they have already chosen, guides carry more of the weight and comparison pages will sit unused.

Second, build for extraction, not for reading time. Lead every page with a verdict in the first 120 words, keep sections to one idea under a literal heading, carry the core claim in a table or list, and source every specific claim. That structural discipline was worth a 2.6x citation lift on our matched set, and it applies to pages you have already published. It is the highest-return work most content teams are not doing.

Third, plan engine by engine and format by format. Cross-engine agreement is only 0.41, so a page tuned for one engine will miss on the others. Aim comparison content at Perplexity, lean on sourced guides and original research for Claude, and hold ChatGPT and AI Mode to the blended playbook. Match your refresh cadence to decay: keep research and documentation current, and leave durable guides alone. Run that loop against the ten prompts that matter to your pipeline, re-measure in thirty days, and let the citation data, not the content calendar, pick what you build next.

KEY TAKEAWAYComparison content earns the most AI citations, but the win is conditional: it has to be extractable, it has to live where the engine expects it, and it has to be aimed at the specific engine your buyers use. Format is the lever, structure is the multiplier, and engine fit decides whether either one pays off.

Cite this research

Cite this research● LIVE
Something Inc. (2026). Which Content Formats AI Engines Cite Most.
6,200 cited URLs across four generative engines, Q2 2026.
Classified by content format and domain type; inter-rater agreement 0.88.
somethinginc.com/insights

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

JB
Josh BernsteinMANAGING PARTNER, SOMETHING INC.

Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.