For most of the last two years, the operating assumption behind generative search measurement has been that the surface is uniform. Everyone asking the same question gets substantially the same answer, so if you can measure that answer once you have measured the market. It was never perfectly true, because personalization and location have always introduced variance, but it was true enough to build reporting on.
That assumption broke quietly in the middle of August, and almost nobody adjusted.
On August 16, Google's Robby Stein and Rajan Patel confirmed that Search now uses Gemini 3.7 Flash in AI Mode, as Barry Schwartz reported at Search Engine Roundtable. The model had been introduced a day earlier as a workhorse model for coding and agents, with no mention of Search. Access is through a picker: click the plus icon, select the model, and search. Availability is limited to Google AI Pro and Ultra subscribers, in English first, with a wider rollout following.
What changed, precisely
Be careful about what this is and is not, because the interesting part is structural rather than dramatic.
It is not a claim that answers got better in a way you can exploit. Stein said the model is better at following instructions and understanding intent. Glenn Gabe, testing it early, said he was initially not seeing a large difference, and later found specific cases where the newer model read intent better than the default. That is an honest read of an incremental improvement, not a step change, and anyone selling you a tactic based on it is ahead of the evidence.
What did change is the shape of the surface. AI Mode acquired a user-selectable variable that affects which sources get pulled into an answer, and that variable is gated by payment. Before August, the question was whether you appear in AI Mode. After August, the honest question is which AI Mode, for whom.
“A surface with a model picker is not one surface with some variance. It is several surfaces sharing a name, and the difference is who is standing in front of it.”
This compounds a split that already existed. Google has been pushing AI Mode into AI Overviews, so some users reach AI Mode responses without ever choosing AI Mode. We wrote about that behavior change in the AI Mode and AI Overviews shift, and the model picker adds a second axis of variation on top of the first.
Why AI Mode visibility now depends on who is searching
The reason this matters commercially is that model access correlates with the buyer you most want.
Think about who pays for Google AI Pro or Ultra. Disproportionately, people who use AI tools heavily for work: engineers, analysts, marketers, founders, researchers. In most B2B categories that population overlaps substantially with the people who evaluate and recommend software. They are not a random sample of searchers. They are closer to a sample of your buying committee.
So the configuration you are least likely to be measuring, because it requires a paid account to observe, is the configuration most likely to be running the queries that decide deals. That is an uncomfortable inversion and it is worth sitting with before deciding it does not apply to you.
Three configurations, three different questions
Rather than a single AI Mode number, treat the surface as three observable states, each answering a different question. The table below is the framing we now use in client reporting.
| CONFIGURATION | WHO SEES IT | WHAT MEASURING IT TELLS YOU | OBSERVATION COST |
|---|---|---|---|
| Default AI Mode | Everyone, including free and logged-out users | Baseline reach across the widest population, and the only state most tooling can see | Low, this is what current vendors measure |
| AI Mode with the newer model selected | Google AI Pro and Ultra subscribers who use the picker | Behavior in front of a heavy AI-using population that skews toward technical buyers | Requires a paid account and manual or bespoke collection |
| AI Mode reached through an expanded AI Overview | Users who never chose AI Mode at all | Incidental exposure, the largest volume state and the least intentional | Medium, hard to separate from AI Overview measurement |
| AI Overviews proper | Default Search users on qualifying queries | A related but distinct surface, only 13.7% URL overlap with AI Mode per Ahrefs | Low, well covered by existing tooling |
Notice the last row. Ahrefs measured URL overlap between AI Overviews and AI Mode at 13.7% in December 2025, and SE Ranking put it at 10.7% of URLs and 16% of domains in August 2025. Those figures were the original argument for treating AI Overviews and AI Mode as separate surfaces. The model picker applies the same logic one level deeper, inside AI Mode itself, and there is no published overlap figure for the model states yet because the split is three weeks old.
That absence is the honest position. Nobody currently knows how much source selection diverges between models in AI Mode. Anyone who tells you they do is extrapolating from a handful of prompts.
It is also worth noting that AI Mode has been gaining presentation formats independently of the model change, including the link carousel treatment on developing stories. Format and model are separate variables that move on separate schedules, and a visibility shift you observe this quarter could be either one. Recording both alongside each result is the only way to tell them apart later, and it costs one extra field.
Google's own reporting merges what the product is splitting
Here is the part that makes this a reporting problem rather than an academic one.
Google's Search generative AI performance report went fully global on August 31 and is now the authoritative first-party source for generative visibility. It covers AI Overviews, AI Mode and AI Overviews in Discover. It reports impressions by page, country, device and date. It does not report clicks, and it does not decompose the surfaces.
So at exactly the moment AI Mode fragments into multiple configurations, the best first-party data available rolls all generative surfaces into a single impression count. Your Google-sourced number gets less granular while the underlying product gets more varied. Those two trends running in opposite directions is the defining measurement condition of this quarter, and it is why we published the error bar model for AI search measurement rather than another set of benchmarks.
How much of the AI Mode question each observation method can currently answer, given the model picker split (Something Inc. assessment)
The practical consequence is that the phrase AI Mode visibility, used without qualification, has stopped carrying information. It needs a modifier now, in the same way that ranking needed one once results became personalized and localized.
The decision: what to measure until this settles
This is a decision under uncertainty, and there are three defensible positions. Only one of them is right for most enterprise programs.
| POSITION | THE CASE FOR IT | THE CASE AGAINST IT | FITS YOU IF |
|---|---|---|---|
| Keep measuring the default only | Cheapest, comparable to existing history, covers the largest population | Systematically blind to the configuration your best buyers may use | Your category skews toward non-technical buyers and low AI tool adoption |
| Add manual sampling on a paid account | Directly observes the state you cannot otherwise see, small fixed cost | Small samples, nondeterministic answers, no automation available yet | Your buyers are technical and a handful of queries decide large deals |
| Wait for tooling to catch up | Avoids building something vendors will ship in two quarters | You will have no baseline from the period when the split appeared | You have no queries where the answer materially changes a purchase |
For most enterprise B2B programs the middle position is correct, and it is cheaper than it sounds. You do not need to sample your whole prompt set on a paid account. You need to sample the ten or twenty queries that actually decide deals, on a schedule, with the configuration recorded next to each result. That is a person for an hour a week, and it produces the only data anyone will have from this period.
What you should not do is quietly change what your existing AI Mode metric measures. If your vendor starts observing a different configuration, your trend line breaks and the break will not be labeled. Ask them, in writing, which configuration their AI Mode data reflects and whether that has changed since July. It is a fair question and the answer belongs in your documentation either way.
What to do this month
Four things, none of them large.
Add a configuration field to how you record AI Mode observations. Whether that is a column in a sheet or a dimension in a platform, every AI Mode data point from here should carry which state produced it. Retrofitting this later is impossible, because the information was never captured.
Get one paid account and start a manual sample on your highest value queries this week. Twenty queries, recorded weekly, with the model selection noted. In six months that file will be the most valuable measurement asset you own, precisely because almost nobody started one.
Stop reporting a single AI Mode figure to leadership without a modifier. Say default AI Mode, or say sampled on the subscriber configuration. The extra three words prevent a conversation in Q4 that you will otherwise lose, and they cost nothing now.
And keep the underlying work pointed where the evidence still supports it. Nothing about the model picker changes what earns a citation. Clear structure, demonstrated authority and machine access were the requirements before August and they are the requirements now, which is why the generative engine optimization roadmap should not be rewritten around a picker. What changes is the reporting layer, and specifically the confidence you are entitled to attach to a number. For B2B SaaS teams whose buyers live in AI tools, that confidence is currently lower than your dashboard implies, and the fix starts with admitting which configuration you have been watching all along.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.