Something Inc.LoginSchedule a free consultation
GEO

Everyone is calling it AI Overviews now. That is a trap.

Bing is testing Google's label on its own answer box. Four different retrieval systems are converging on one name, and the source sets behind them barely overlap.

GEOREPORTED AUG 27, 2026
13.7%
source overlap between Google's AI Overviews and AI Mode, despite agreeing on the answer 86% of the time
28.3%
of ChatGPT's most-cited pages that have zero Google organic visibility
45%
of surveyed B2B buyers using Copilot for product research, against 71% for ChatGPT
THE SHORT VERSIONBing is testing renaming its Copilot Search answer to AI Overview, matching Google's terminology. The naming is converging. The retrieval is not. If your reporting collapses four engines into one AI Overviews number because they now share a label, you will lose the only signal that tells you where to spend.

Naming is never only naming. When two competing products adopt the same word for the same slot on the page, buyers stop distinguishing between them, and shortly afterwards so do the teams optimizing for them. That is the risk in a small test Bing ran this week.

Bing is borrowing Google's label

Sachin Patel spotted Bing testing a renamed answer box on mobile, swapping the Copilot Search label for AI Overview while keeping the Copilot logo in place, and shared it via X and Search Engine Watch before Search Engine Roundtable picked it up. Microsoft has not commented. It may not ship.

Treat it as a signal rather than an event. Microsoft spent two years building Copilot as a brand and attaching it to everything, and a test that subordinates that brand to a competitor's generic-sounding label is a meaningful concession. The most plausible read is that Copilot Search was not communicating what the box was, while AI Overview does, because Google's version has been in front of far more people for far longer. Microsoft is trading brand differentiation for immediate comprehension.

The interesting question is not whether Bing ships it. It is what happens to everyone's measurement when four different systems all present something called an AI overview.

There is precedent for how this goes. Search marketing has repeatedly adopted a dominant player's vocabulary and then quietly assumed the underlying mechanics matched the name. Featured snippets, rich results and knowledge panels all acquired generic usage that flattened real differences between engines, and in each case the flattening showed up first in reporting and only later in strategy. The pattern is consistent: the label spreads faster than the understanding, and by the time anyone checks, the metric built on the label has been in a board deck for two quarters.

Why a shared name is a measurement problem

Shared vocabulary produces shared metrics, and shared metrics produce averaging. The moment your dashboard has a column labelled AI Overviews rather than four columns labelled by engine, someone will average them, and the average will be reported upward as a single visibility number.

That number would be close to meaningless, and we can show why with data from inside a single company. Ahrefs analysis of over a billion data points found that Google's AI Overviews and Google's AI Mode reach the same conclusion to a query 86% of the time, while citing different sources 86.3% of the time. The source overlap between them is 13.7%. We unpacked that finding when we looked at where AI Mode and AI Overviews actually source from, and it remains the single most clarifying statistic in this field.

Two surfaces owned by the same company, answering the same question the same way, share one source in seven. A shared label across four companies will hide more than that, not less.

Sit with the implication. These two surfaces share infrastructure, an index, a corpus and a corporate owner. They agree on what the right answer is. And they still pull from almost entirely different sets of pages to support it. If that is the divergence between siblings, the divergence between Google, Microsoft, OpenAI and Perplexity is not something a single averaged metric can survive.

The engines are not converging where it counts

Convergence is real at the interface layer and largely absent underneath it. It is worth being precise about which layers are actually moving together, because the answer determines how much of your work transfers between engines.

LAYERCONVERGING?WHAT THIS MEANS FOR YOU
Name and placementYesUsers will treat all four as the same feature, and so will internal stakeholders
Interaction patternMostlySummary, then citations, then follow-up. Optimization for extractability transfers well
Retrieval corpusNoDifferent indexes and different live-fetch behavior produce different candidate sets
Source selectionNo13.7% overlap between two Google surfaces alone; cross-engine overlap is lower
Citation displayNoLink carousels, inline chips and footnote lists reward different page structures
Measurement accessNoEach engine exposes different data, and most expose almost none

The top two rows are why a single label is seductive. The bottom four are why acting on it is expensive. Optimization work that targets the interaction pattern, clear structure, self-contained claims, direct answers near the top, genuinely does transfer across every engine, and that is the portion of GEO work worth doing once. Everything below that line is engine-specific and does not transfer, and treating it as though it does is how teams end up optimizing hard for a surface their buyers do not use.

The corpus divergence is the most consequential and the least visible. Our analysis of ChatGPT's citation behavior found that 28.3% of its most-cited pages have zero Google organic visibility, which we covered in the gap between ChatGPT citations and Google rankings. A page can be invisible in Google and heavily cited in ChatGPT. No amount of shared labelling changes that, and no averaged metric will ever reveal it.

How to compare AI Overviews across engines properly

The fix is not complicated, it is just unglamorous, and it requires refusing a simplification that stakeholders will actively ask you for. Three rules hold up well in practice.

1Never average across enginesReport per engine, always, even when the label is identical. If an executive summary needs one number, make it the weighted figure for the engines your buyers actually use, and state the weighting next to it. An unweighted average across four engines describes no real user.
2Weight by your buyer mix, not by engine sizeThe Semrush survey of 519 US B2B decision makers put ChatGPT at 71% for product research, Gemini at 61% and Copilot at 45%. Those are population figures. Your category may skew heavily one way, and a Microsoft-heavy enterprise buyer base makes Copilot performance matter far more to you than its overall share suggests.
3Track source overlap, not just presencePresence tells you that you appeared. Overlap tells you whether the pages earning your citations are the same pages across engines, which is what determines whether a content investment pays once or four times. Low overlap is not failure, it is a budgeting input.
ChatGPT71%
Google Gemini61%
Microsoft Copilot45%

Tools used for product research by US B2B decision makers. Semrush survey, 519 valid respondents, March to April 2026.

Note what that chart does not say. It does not say Copilot is 45% as important as ChatGPT to your business, because these are population figures across all US B2B professionals and your buyers are not a random sample of them. An enterprise selling into large Microsoft-centric organizations, where Copilot is deployed by default and sits inside the tools people already have open, will see a distribution that looks nothing like this. So will a developer-tools company whose buyers live in a different set of habits entirely. Use published figures as a starting hypothesis and then check it against your own referral data and your own sales conversations, because the weighting is the entire point and getting it wrong sends your budget to the wrong engine with full confidence.

The third rule is the one that changes behavior. Teams tracking presence alone see four numbers moving semi-independently and cannot explain why. Teams tracking which pages earn the citations discover that a single asset is carrying them in one engine while a different asset carries them in another, and that is directly actionable in a way a presence percentage never is.

What to actually optimize when the surfaces differ

Split the work into the portion that transfers and the portion that does not, and fund them differently. This is the practical consequence of everything above.

The transferable portion is structural and deserves the majority of the budget, because it pays into every engine at once and into classic search as well. Self-contained claims that survive being lifted out of context. Direct answers positioned near the top rather than after four paragraphs of preamble. Real tables instead of prose describing what a table would say. Named authors and visible dates. Clean machine access. None of this is engine-specific and none of it is new, which is exactly why it is reliable.

The non-transferable portion is corpus work: getting represented on the specific sources a specific engine leans on. This is where per-engine divergence bites, where the cost is highest, and where you should be most disciplined about only doing it for engines that matter to your buyers. Do it for one or two engines properly rather than four badly.

One caution on tooling. Vendors will move quickly to offer a unified AI Overviews metric across engines, because the shared label makes it saleable and because customers will ask for it. Interrogate what any such number is actually measuring before it enters a board deck, since a metric that spans engines with single-digit source overlap is an aggregation, not a measurement. This is the same category of dependency risk we set out in the search dependency audit framework: a number you cannot decompose is a number you cannot act on.

Do this before your next AI visibility report

Open your current AI visibility reporting and answer one question: can you tell, per engine, which of your pages earned citations last month? If the answer is no, that is the gap to close before anything else, and it is a bigger constraint on your GEO program than any individual optimization you might run this quarter.

Then add a source-overlap line: of the pages earning citations, how many earn them in more than one engine. Watch that number over a quarter. If it is high, your structural work is doing the heavy lifting and you should keep funding it. If it is low, you are running four separate visibility programs whether or not you have budgeted for four, and you should decide deliberately which ones you are actually committing to.

Bing adopting Google's vocabulary is a small thing that makes a large error easier to make. The label is converging because it helps users understand what they are looking at, which is legitimate. The systems behind it are diverging because they are built on different indexes with different retrieval and different incentives, and that is not changing soon. Categories where the buyers are themselves technical, AI and machine learning especially, tend to spread across engines more than the population averages suggest, which makes per-engine measurement more valuable there rather than less. Keep the columns separate, weight by the buyers you actually have, and treat any single cross-engine number with the suspicion it has earned. That discipline is most of what separates a working generative engine optimization program from a dashboard that reassures everyone and directs nothing.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

JB
Josh BernsteinMANAGING PARTNER, SOMETHING INC.

Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.