Google started serving AI Mode answers with Gemini 3.8 Flash on September 2, one model generation and a few weeks after 3.7 Flash. Within twenty four hours, people who watch these surfaces closely were posting screenshots of answers that carried no source links at all, mostly on broad top-of-funnel queries. Robby Stein, VP of Product for Google Search, replied in public that it was not working as intended and that a fix was coming. So the citations were not lost to a competitor, or to a content quality problem, or to anything anyone did on their own site. They were lost to a deployment.
That is the part worth sitting with. Every serious brand now has some version of a citation dashboard, and most of those dashboards would have shown a cliff on September 3. Almost none of them would have shown the reason, because the reason was not in the data they collect. Search Engine Land reported the fix commitment the same evening, after Gagan Ghotra and Glenn Gabe had posted the screenshots and Barry Schwartz wrote it up. The whole story lived on a social feed and two trade sites for about a day.
What happened to AI Mode citations on September 2
The sequence is short and it is worth writing down precisely, because the precision is the whole value. Google announced that Gemini 3.8 Flash had landed in Search, available in AI Mode for Google AI Pro and Ultra subscribers, selectable from the model dropdown. Google DeepMind describes 3.8 Flash as its most capable workhorse model to date, with gains over 3.7 Flash on software engineering, agentic tasks and multi-step reasoning. None of those gains are about citation behavior, and citation behavior is not usually part of a model card.
A day later the screenshots started. Answers generated on the new model, particularly on broad informational queries, came back as prose with nothing attached. No source chips, no links, nothing for a publisher to be measured by. Stein's response was direct and quick: this is not intended, a fix is coming. That is the correct answer from Google and it should be read as good news. It is also, from where you sit, completely outside your control.
Notice what the exposure looked like. This was scoped to a model that paying subscribers could pick. It was not a ranking change, not a spam action, not a core update. It had no announcement, no rollout note, and no entry in any of the places a search team habitually watches. If your team had spent September 3 debugging a citation drop by auditing schema, checking crawler access and rewriting a comparison page, every hour of that would have been wasted on a problem that resolved itself when Google shipped the fix.
This is not a one-off. It is what optimizing against a product on a fast release cadence feels like from the outside. The behavior differences between surfaces are already large enough that we treat them as separate channels, which is the argument in how AI Mode and AI Overviews actually behave differently. Add per-model differences inside a single surface and the number of distinct answer environments you are being measured across stops being small.
Supply and share are two different numbers
Most citation reporting collapses two questions into one metric. The first question is whether the answer surface emitted any citations at all for a given query. Call that citation supply. The second is whether you were among them. Call that citation share. When supply is healthy and your share drops, that is a content, authority or competitive problem and it belongs to your team. When supply itself goes to zero, nothing about your site is being measured, and the number in your dashboard is not a performance number. It is a status number for somebody else's product.
| WHAT YOU OBSERVE | SUPPLY | SHARE | WHO OWNS THE FIX |
|---|---|---|---|
| Answer returns three sources, none of them yours | Healthy | Zero for you | You. This is a content and authority problem |
| Answer returns no sources for anyone on the query set | Zero | Undefined | The engine. Wait for the fix and annotate the chart |
| Answer returns sources on one model, none on another | Model-dependent | Model-dependent | Split the report by model before drawing a conclusion |
| Answer cites you but sends no measurable traffic | Healthy | Healthy | You, but it is an attribution problem, not a visibility one |
| Answer cites a syndicated copy of your page instead of you | Healthy | Misattributed | You. Canonical, licensing and distribution question |
The row that gets missed is the second one. A reporting pipeline that only tracks your own mention rate cannot distinguish it from row one, and rows one and two lead to opposite decisions. This is the same reason we push clients toward reporting that carries uncertainty explicitly instead of a single clean line, which is the case we made in the error bars framework for AI search reporting.
Fixing this costs almost nothing. Add a small control set of queries to whatever you already run: ten to twenty broad, non-branded, top-of-funnel questions in your category that a healthy answer surface should always cite something for. You do not care who gets cited on them. You care only whether anyone does. When your brand's number drops and the control set is still emitting sources, the problem is yours. When the control set goes quiet at the same moment, the problem is upstream and your job is to annotate, not to act.
Illustrative triage weighting: where the cause usually sits when a citation number drops overnight, based on how we route incidents in client programs
Treat that split as a working prior, not a finding. The point of it is directional: the single most likely explanation for a sudden overnight change is not something you did, and the cheapest first check is the one that rules that out.
The 72 hour freeze, and what belongs inside it
The operational rule we hold teams to is simple. A citation drop does not become a work ticket for 72 hours. It becomes an investigation. Those are different things and conflating them is what burns a quarter of your content capacity on noise.
The reason to formalize this is not tidiness. It is that a drop with no explanation attracts explanations, and the explanations that get attached in the first 24 hours tend to be whatever the loudest person in the meeting believes about AI search. A written freeze protocol gives you a legitimate answer to give an executive on day one that is not a guess. Most of the credibility damage from these events comes from that first meeting, not from the drop.
“A number that moved because someone else shipped is not a result. It is weather. Report it as weather.”
Build a model release log next to your AI Mode citations log
The missing input is boring and it is a spreadsheet. Alongside whatever you use to track AI Mode citations, keep a dated log of the surface itself: which model is serving, when it changed, what the vendor said about it, and which subscriber tier or region was affected. Twelve months of that log is the difference between a chart you can defend and a chart you can only apologize for.
None of this is exotic. It is change management applied to a channel that has not had any, and it is the kind of instrumentation work we set up before we touch content in a generative engine optimization engagement, because the alternative is optimizing blind against a surface that redefines itself every few weeks. The same discipline sits underneath the dashboards our reporting and analytics team builds, where an annotation layer is a first-class object and not something bolted on after the first bad quarter.
There is a second-order benefit. Once you have a release log, you can start answering questions that are otherwise unanswerable, like whether your citation rate is stable across model versions or whether it quietly depends on one of them. That is a real strategic finding. A brand whose visibility only holds on the older model has a fragility problem it needs to know about now, not the next time a version ships.
What to do this week
Start with the control set, because it takes an afternoon and it is the check that pays for itself the first time something moves. Ten to twenty broad category questions, run across the surfaces that matter to you, logged weekly, measuring only whether sources appear at all. Nothing about your brand in the query, nothing about your brand in the measurement.
Then write the freeze rule down and get someone senior to agree to it before you need it. A 72 hour investigation window is easy to agree to in a calm week and impossible to introduce in the middle of a drop. Put it in the same document as your escalation path, and name who is allowed to declare an incident over and who is allowed to authorize content work off the back of one.
Start the release log with today's entry and backfill the ones you can source. Gemini 3.8 Flash entering AI Mode on September 2, with the citation behavior reported and acknowledged on September 3, is a legitimate first row. Add the model picker tiering while you are in there, since the availability of different models to different subscribers is itself a segmentation of your addressable audience, which is the argument in what the AI Mode model picker does to visibility.
Then change one line in your next report. Where you currently write that citations fell, write what supply did and what share did, separately, and say which one moved. That single change will do more for how your work is understood inside the business than another quarter of content will. The teams that lose budget in this channel almost never lose it because the work was bad. They lose it because a chart moved and nobody in the room could say why.
Google will fix the citation gap, probably before most brands notice it happened. That is not the point. The point is that it could happen at all, without warning, on a surface that a lot of companies are now betting a meaningful share of their pipeline on. Build the log. The next one will not come with a public acknowledgement inside 24 hours, and you will want the receipts.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.