Amazon, Walmart, and Wikipedia are not competitors in your category. They are the landlords. Ahrefs' latest citation dataset puts a number on something most GEO teams have felt without naming: a handful of mega-authority domains absorb a disproportionate share of AI citations before your brand ever gets a turn at the podium.
Every GEO plan we review starts from the same unstated assumption: publish enough good pages and the citations follow. That assumption holds for a real slice of queries. It falls apart the moment the query already has an answer sitting on a domain with more authority than any brand will build in a decade. AI citations are not evenly distributed across the web, and pretending they are is one of the more expensive mistakes we see in year-one generative engine optimization budgets.
This isn't a niche problem that only affects small brands chasing broad terms. It affects every content team that treats the roadmap as a flat list of keywords to cover, one page per gap, without first asking who already answers that gap and how well they answer it. A well-run editorial calendar and a well-run GEO program look similar on a whiteboard. They diverge the moment you ask which rows on that calendar are competing against Wikipedia's reference authority or Amazon's product-data depth, and which ones are genuinely open. Most calendars don't ask, because the tooling that produces them was built for classic search, where a page with better on-page optimization and more backlinks could eventually out-rank almost anything given enough time and budget. Generative engine citation doesn't reward patience the same way. It rewards whichever source the model already trusts enough to lift verbatim, and for a large share of queries, that source was decided before your brand's roadmap existed.
Why AI citations concentrate on three domains
Ahrefs tracks this directly. Its ongoing analysis of the 50 most-cited websites in Copilot, built with Ahrefs' Agent A tool and refreshed monthly against a June 2026 base, ranks the domains Microsoft Copilot cites most often across a large sample of tracked answers. The top of that list is not a long tail spread thin across thousands of sites. It's three familiar names doing an outsized share of the work.
| DOMAIN | SHARE OF COPILOT'S MOST-CITED CITATIONS |
|---|---|
| Amazon | 14.6% |
| Walmart | 10.2% |
| Wikipedia | 9.6% |
| Combined (top 3) | 34.4% |
Add those three rows and more than a third of the tracked citation share flows to domains that were never optimizing for your keyword. Amazon and Walmart earn their share through raw retail authority and product-data depth built over two decades. Wikipedia earns its through reference authority Google spent just as long reinforcing, authority that AI engines inherited wholesale the moment they started citing sources instead of just ranking links. None of the three built that position for your category. They built it once, broadly, and now collect citation share across thousands of categories at the same time, yours included, without spending a dollar on your keyword specifically.
It's worth being precise about what the Ahrefs dataset is and isn't measuring, since the distinction matters for how you act on it. It's a domain-level view of who Copilot cites most often, built with Ahrefs' Agent A tool and updated monthly against a rolling base, currently anchored to June 2026. It doesn't tell you that Amazon or Wikipedia wins every query in every category. It tells you that across a large, broad sample of tracked answers, those three domains show up in the citation pool often enough to account for over a third of total share. Some of that is retail-intent queries where Amazon and Walmart are structurally correct answers. Some of it is reference and definitional queries where Wikipedia is the obvious source. The takeaway isn't that these three domains are unbeatable everywhere. It's that they've already won a specific, identifiable set of query types, and no amount of owned-page volume changes who wins those types.
What Something Inc.'s own AI citation data adds
Ahrefs' dataset measures who gets cited across the open web, at the domain level. Our own research measures something adjacent: how often a single well-optimized page gets picked up once it exists, tracked across a few thousand high-intent B2B prompts on ChatGPT, Perplexity, Claude, AI Mode, and Gemini. The two numbers are not contradictory. They describe two different bottlenecks in the same funnel, and most GEO reporting only ever measures the second one.
That gap in what gets measured is the quiet reason so many GEO programs feel like they're doing everything right while citation share stays flat. A team ships an llms.txt file, tightens its schema, publishes a well-structured comparison page, and watches mention rate climb on the queries they're tracking directly. All real progress. None of it moves the needle on the query types where a domain outside the brand's control already sits at the top of the citation pool before the model even starts weighing which brand page to surface. Mention rate answers were you picked. It doesn't answer were you even in the running. Programs that only track the first number can look successful in the dashboard while losing the larger share of the category's total answer volume to Amazon, Walmart, or Wikipedia without ever noticing.
Mention rate for a well-optimized page, by engine (Something Inc. research)
A mention rate in the high 60s and 70s looks like a page that's winning, and conditionally, it is. That number describes how often the page gets picked up once an engine has already decided your category's answer includes brand-level pages at all. Ahrefs' data describes the earlier, harder bottleneck: whether the engine reaches for a brand page in the first place, or whether it reaches for Amazon, Walmart, or Wikipedia before your page is ever considered as a candidate. Build a page that mentions correctly three out of four times and it still won't surface if the query resolves to a domain sitting two or three tiers of authority above you before the shortlist gets made.
Comparison and alternatives content earns more AI citations than any other format we've measured, 32.5% of everything we tracked, by a wide margin over how-to guides, product docs, and raw research. That's the strongest argument for building where you can genuinely win. It's also, by the same logic, the format Amazon and Wikipedia are least likely to preempt, since neither is in the business of writing a head-to-head comparison of your category's vendors. That's the build side of the ledger. The earn side is everything else: the generic, already-answered, top-of-funnel query where a reference domain or a retail giant got there first and isn't moving. Earning presence on those domains instead of competing with them is the same logic behind the case for third-party citations over owned pages, and it sits next to a related finding worth knowing: 82% of AI citations trace back to earned media rather than brand-owned content, which we cover in more depth here.
Build vs. earn: the decision most GEO programs never make explicitly
Almost every GEO program we take over already has an owned-content roadmap. Almost none of them has a documented answer to a simpler, earlier question: for this specific content gap, are we trying to out-rank the domain that already owns the answer, or are we trying to get named on it? The default answer, chosen by omission rather than decision, is always build. Content teams know how to brief, write, and publish a page. Most of them don't have a standing process for getting a data point into a Wikipedia citation or a structured feed onto Amazon, so that option never even makes it onto the roadmap to be weighed and rejected. It just never comes up as a choice.
There's a structural reason the build option always wins by default, and it isn't strategy, it's org chart. A content team can commission, edit, and publish an owned page entirely inside its own workflow, on its own timeline, with a single approval chain. Earning a listing on Amazon's structured feed, a citable data point on Wikipedia, or a contributor byline on a publication the engines already trust requires a different skill set, usually one the content team doesn't have and isn't measured against. So the roadmap defaults to what the team already knows how to execute, and the domains that actually dominate the category's citation pool never enter the conversation. That's not a strategic bet against earning. It's the absence of a bet at all, and the two look identical on a roadmap until you check which queries the brand is actually showing up on six months later.
A decision table for your content gaps
Run this against your next content gap list before a single brief gets written. For each gap, name the domain that currently owns the answer, then make the build-or-earn call on purpose instead of by default.
| CONTENT GAP | WHO ALREADY OWNS THE ANSWER | BUILD OR EARN | THE MOVE |
|---|---|---|---|
| Generic category query ("best project management software") | Wikipedia, Amazon-adjacent marketplaces, aggregator sites | Earn | Get your data into the sources those pages already cite; a new owned page rarely displaces the incumbent |
| Head-to-head vendor comparison ("X vs. Y") | Usually nobody, or a thin affiliate page | Build | Ship the comparison yourself. It's the format that earns the most AI citations of any we've measured |
| Definitional or glossary query ("what is generative engine optimization") | Wikipedia and reference sites | Earn | Contribute a citable, sourced fact a reference editor will actually use, not a competing definition page |
| Direct purchase query ("buy [product] online") | Amazon, Walmart | Earn | Structured product feeds and marketplace listings beat a landing page for this query type |
| Practitioner how-to or troubleshooting query | Reddit and niche forums | Earn | Real, attributed participation in the threads already ranking beats a new owned guide |
| Your category's evaluation criteria ("how to choose a [category] vendor") | Usually open | Build | Nobody else can publish your point of view; this is the highest-leverage owned content on the list |
Notice the pattern. The build column is short and specific: comparisons, and your own evaluation criteria, the content only you can credibly write. Everything else, the generic category term, the definitional query, the direct purchase intent, the practitioner troubleshooting thread, belongs to a domain that already won that fight before your brand existed. Digital PR and earned placement work is how you show up on those domains without pretending you can out-publish them, and it's the same operating model we rebuilt for a B2B staffing platform that had spent two years shipping owned pages into a category where the AI answer was never going to cite it for them.
Do this next
Pull your current content gap list and run every row through the table above before you brief a single new page. Where the answer is earn, that work routes through structured data, marketplace presence, and earned placement, not a content calendar. Where the answer is build, build with intent: lead with a verdict, cite every claim, and make it the kind of comparison an engine can lift without rewriting it first.
“The brands winning AI citations aren't the ones publishing the most pages. They're the ones who stopped fighting Amazon and Wikipedia for queries those domains already own, and started showing up on them instead.”
Run the audit this week. For every open content gap, write down which domain currently owns the citation, and put a build or earn label next to it before anyone opens a doc. That one-line decision, made on purpose instead of by accident, is the difference between a GEO budget that compounds and one that spends a year publishing pages Amazon was always going to outrank.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.