Ask ChatGPT, Perplexity, and Gemini the exact same question and you'd expect some disagreement about which sources deserve the answer. What a Manchester-based GEO agency found this week is closer to three separate answer sheets: 30 identical buyer questions, 104 different sources cited, and only three businesses that all three engines agreed on. If your GEO plan assumes one page can win all three engines, this is the study that says otherwise.
What the study actually tested
Alex Golombeck, founder of The Happy Cat, ran the test on August 19 and 20, 2026: identical buyer-intent questions, the kind a real prospect types before hiring someone, put to ChatGPT, Perplexity, and Gemini. Queries included "best AI visibility agency UK" and "AI SEO agency Manchester," phrased the way a buyer actually searches rather than the way a marketer would optimize for. Thirty answers came back, and every citation in every answer was logged, screenshotted, and published.
The methodology detail that makes this worth taking seriously is the one most vendor studies skip: The Happy Cat published its own result. It appeared in zero of the 30 answers, and Golombeck committed publicly to re-running the identical 30 queries on September 16 and October 14, 2026, so the baseline is checkable against a moving target instead of a single flattering snapshot.
That willingness to publish an unflattering baseline is rare in this category. Most GEO case studies show up only after the win, with the before number reconstructed from memory or left out entirely. A study that starts by admitting the agency running it appeared nowhere, and commits to re-checking on a fixed public calendar, is built to be falsified rather than just believed. That's the standard worth holding every AI citation claim to, including the ones in this article.
Why a 30-query study is still worth taking seriously
Thirty questions is a small sample next to a dataset like Semrush's 126-million-prompt AI Visibility Index, and it would be fair to wonder whether 104 sources across 30 answers is just noise. Two things argue against that. First, the finding isn't really the raw number, it's the shape: a 3-in-30 unanimous rate and an 11-in-104 repeat rate both describe the same pattern, that engines mostly disagree rather than mostly agree, and a pattern that shows up at both a small, controlled scale and a large, aggregated scale is more trustworthy than either alone. Second, the study is narrow by design. General "best X" queries pull from a much wider source pool than a single company's branded terms would, which makes disagreement more likely to begin with, but that's also the exact kind of query that decides whether a prospective buyer ever hears your name before they've hit your website.
104 sources, 11 repeats, three unanimous picks
The headline number is the spread. Thirty answers, if the engines mostly agreed on who deserves to be cited, should produce a fairly short, overlapping source list. Instead they produced 104 distinct sources. Only 11 of those got cited more than once across the full set, and just 3 businesses showed up in the answers from all three engines at the same time.
“The practical upshot is that AI visibility is currently a lottery with three different sets of rules.”
That's Golombeck's own read on the data, and the numbers back it up. A brand that nails ChatGPT's citation preferences and ignores Perplexity's is optimizing for roughly a third of the actual answer surface. The old assumption in GEO, that one well-structured page earns citations everywhere once it clears a quality bar, doesn't survive contact with a 104-source spread from 30 questions.
It's worth sitting with what "three unanimous picks" out of 30 questions actually implies for a category with real competition. If the three engines drew from a shared, stable ranking of trustworthy sources the way classic Google search results roughly do, you'd expect the same handful of names to keep surfacing across most of the questions. Instead, 93 of the 104 sources appeared only once, meaning the overwhelming majority of citations were one engine's individual judgment call, not a consensus pick multiple systems converged on independently. That's a fundamentally different game than ranking for a keyword, where the target is one ordered list. Here the target is three separate, only loosely correlated lists, and a brand can lead one and be invisible on the other two at the same time, which is a harder problem to solve and a harder problem to even notice without checking each engine on its own.
Three engines, three different citation habits
The more useful part of the study isn't the overlap number. It's what each engine was actually doing when it picked a source, because the three showed distinct, repeatable habits rather than random noise.
| ENGINE | WHAT IT FAVORED | CITATION HABIT |
|---|---|---|
| ChatGPT | Businesses' own websites | Every source it cited carried a utm_source=chatgpt.com tracking tag |
| Perplexity | Third-party comparisons and roundups | Almost never cited a business's own site; pulled from industry roundups and, in one case, a press release |
| Gemini | Long-form published articles | Quoted specific passages verbatim, rewarding businesses that had written something substantive enough to lift directly |
This matches what we see in our own citation data
Small-sample studies are easy to dismiss as noise, so it matters that the shape of this finding lines up with a larger dataset. When we ran our own citation study across 4,100 high-intent B2B queries, a source cited by ChatGPT was cited by Perplexity only 44% of the time. That's a different query set, a different sample size, and a different month, and it points at the same conclusion: engine agreement is closer to a coin flip than a consensus.
Where a well-optimized page gets cited, by engine (mention rate, Something Inc. citation data)
Semrush's 2026 AI Visibility Index, built from 126 million AI search prompts, adds another angle on the same gap: ChatGPT cites an average of 15 sources per response, while Gemini cites around 3. An engine pulling from 15 sources has more room to include a new or smaller brand than one pulling from 3, which is part of why comparison content that names multiple options tends to do better on ChatGPT specifically than on Gemini, where the bar for the single quotable passage is much higher.
What to actually do about it
In practice that means treating generative engine optimization as three overlapping but distinct jobs rather than one checklist. Structure your own pages so ChatGPT can send buyers there directly: clear service pages, direct claims, nothing buried behind a contact form. Separately, invest in the third-party presence, reviews, comparison roundups, community threads, that Perplexity actually pulls from instead of your own domain. And write at least a few paragraphs precise and specific enough that Gemini can quote them verbatim rather than paraphrasing something vaguer.
This is exactly the gap a technical and content audit is built to find: which of the three jobs your current content is actually doing, and which one it's quietly skipping. For B2B SaaS companies competing in categories where the buyer's first move is asking an AI engine instead of Googling, that gap is the difference between showing up in one answer out of three and showing up in all of them.
The 104-source spread in this study is a snapshot, not a law of nature, and Golombeck's committed re-runs in September and October will show whether the gap narrows as the engines mature or stays this wide. Until an engine changes how it picks sources, building for only one of the three is building for a third of your buyers at best, and the businesses that showed up in all three answers this week almost certainly weren't optimizing for a single engine when they earned that spot. They more likely had a clean, direct site an assistant was comfortable linking to, a track record with the third parties Perplexity trusts, and content specific enough for Gemini to quote, three separate investments that happened to land in the same 30 questions.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.