Ask an outbound team how many sending domains they run and the number keeps going up. Twenty was a lot two years ago. Forty is common now. The reasoning behind the growth is sound at every individual step: spread the volume, keep each mailbox modest, never let one domain carry enough traffic to look like a mass sender, and if a domain burns, retire it and provision another. It is a sensible response to a filtering regime that punishes volume concentration. It also quietly removed the only instrument that tells you when you are about to be filtered, and almost nobody made that trade on purpose.
The Bulk Sender Rules Assume A Sender You Are Not
Start with what the receivers actually wrote down, because it is short and specific. Google's sender guidelines apply an enhanced set of requirements to senders above 5,000 messages a day to Gmail accounts: aligned SPF and DKIM, a DMARC record on the sending domain, valid reverse DNS, TLS on the connection, and one click unsubscribe on marketing and subscribed mail. The same document gives the number everyone quotes: keep spam rates reported in Postmaster Tools below 0.30%, and Google goes further, telling senders to stay below 0.10% and never reach 0.30% at all.
Microsoft drew the line in the same place. Its high volume requirements cover senders pushing 5,000 or more messages a day at outlook.com, hotmail.com and msn.com, with SPF and DKIM passing and DMARC at p=none or stricter, aligned to at least one of them. Validity documented the enforcement change, which started on May 5, 2025, and noted that Microsoft moved from its original plan of junking non compliant mail to rejecting it outright. Visible, working unsubscribe is required and must be honoured immediately, though the one click header standard is supported rather than mandated.
Read those two policies as an outbound operator and the conclusion writes itself: none of this is aimed at us. A forty domain estate never puts 5,000 messages a day behind any single sending domain. The bulk sender chapter is for the marketing team and their ESP, and outbound sits comfortably below the line. That reading is correct on the letter of the policy and badly wrong on the consequence, because the threshold does two things at once. It decides which senders get held to a written standard, and it decides which senders get told how they are doing.
That second function is the one nobody scoped. Google does not publish a minimum volume for Postmaster Tools reputation data, but the working number deliverability practitioners report is somewhere around 100 messages a day to unique Gmail users before a domain generates anything readable, and materially more than that before the spam rate panel gives you a trend rather than a shrug. A sending domain carrying 120 messages a day across all providers is not clearing that bar for Gmail alone. It is producing a blank panel, and a blank panel is not a clean bill of health.
Cold Email Complaint Rate Has No Volume Floor
Here is the asymmetry that makes the architecture dangerous rather than merely uninformative. The reporting threshold has a volume floor. The consequence does not.
When a recipient clicks report spam, that signal is recorded against the sending domain and the sending IP regardless of whether that domain sent five messages that day or fifty thousand. Filtering decisions are made on reputation, and reputation accrues from user behaviour, authentication results and engagement at whatever volume you happen to be sending. The 0.30% figure is the number Google chose to publish as a requirement for the senders it holds to a written standard. It is not a switch that turns on at 5,001 messages a day and leaves everyone below it unjudged. A domain at 120 sends a day with a bad complaint pattern gets filtered exactly like a domain at 12,000 sends a day with the same pattern. It just gets filtered without ever having seen the number that predicted it.
This is a genuinely different failure from the one we wrote about when Microsoft pulled trap data out of its monitoring feeds. That was visibility the provider took away. This is visibility the sender gave up, in exchange for a volume profile that looks unthreatening. Both leave you flying without instruments, but only one of them is a decision you can revisit this quarter.
Sharding Turns Cold Email Complaint Rate Into Noise
The measurement problem is bad. The statistics problem is worse, and it is the part that practitioners consistently miss because a percentage feels like a percentage no matter what it is computed over.
Take a standard estate and do the arithmetic. The prevailing setup advice across the major sending platforms and outbound agencies is 2 to 3 mailboxes per domain, 30 to 50 messages per mailbox per day once warmed, which puts a domain at roughly 100 to 250 a day. Call it 20 domains, 3 mailboxes each, 40 messages per mailbox: 2,400 messages a day from the estate, 120 a day per domain. Now assume the estate is running well, with a true underlying complaint rate of 0.10%, which is the level Google asks senders to stay under. That is 2.4 complaints a day across the whole estate.
Those 2.4 complaints do not distribute themselves evenly. They land on two or three domains, and on those domains a single complaint against 120 messages reads as 0.83% for the day, nearly three times the ceiling Google publishes. The other seventeen domains read 0.00%. Same estate, same copy, same list, same day. The aggregate number is healthy and every individual number is either perfect or alarming, with nothing in between, because a denominator of 120 cannot express a rate of 0.10% at all. The smallest non zero value that denominator can produce is eight times the target.
Illustrative arithmetic on a 2,400 message a day estate with a true complaint rate of 0.10%, expressed as a share of Google's published 0.30% ceiling and capped at the chart maximum. The scenarios use the prevailing build advice of 2 to 3 mailboxes per domain at 30 to 50 sends each. No provider publishes per domain complaint distributions, so treat these as modelled numbers, not measurements.
Sharding does not reduce the complaint rate. It cannot: the same people receive the same mail and press the same button. What sharding changes is the variance, and it changes it in the direction that hurts. A large denominator absorbs a single bad reaction and reports a number you can act on. A small denominator converts every single bad reaction into a spike that looks like an emergency, and converts every quiet day into a zero that looks like success. Neither reading is useful, and teams end up making infrastructure decisions on whichever one arrived most recently. This is the same measurement trap we described when arguing that a reply rate cannot tell you what is broken, one layer further down the stack.
The Case For Fewer, Bigger Sending Domains
The honest version of this argument has a real cost attached, so state it plainly. Consolidating volume onto fewer domains concentrates risk. If a domain does burn, more of your capacity goes with it, and the whole reason the industry drifted toward sprawl is that a burned domain in a forty domain estate costs you two and a half percent of your sending and a weekend of provisioning. That is a genuine benefit and this is not an argument for throwing it away.
| ESTATE SHAPE | SENDS PER DOMAIN PER DAY | WHAT PER DOMAIN REPUTATION DATA YOU GET | WHAT ONE SPAM COMPLAINT DOES TO THAT DOMAIN'S RATE FOR THE DAY |
|---|---|---|---|
| Heavily sharded, 40 domains at 2 mailboxes | Around 80 | Nothing readable. Gmail volume per domain sits far under the level practitioners report as the floor for reputation data, so the panel stays blank | 1.25%, more than four times Google's published ceiling, on a domain whose true rate may be fine |
| Standard build, 20 domains at 3 mailboxes | Around 120 | Still nothing usable on most domains. Authentication results show, reputation and spam rate generally do not | 0.83%, nearly three times the ceiling. The smallest non zero rate this denominator can express |
| Consolidated, 8 domains at 5 mailboxes | Around 300 | Borderline. Larger domains begin generating reputation signal, enough for a direction of travel rather than a clean trend | 0.33%, just over the ceiling, and a second complaint tells you something rather than doubling an artefact |
| Concentrated, 4 domains at 6 mailboxes | Around 600 | Readable. Enough daily Gmail volume on each domain to populate a spam rate you can trend week over week | 0.17%, inside the ceiling. A rate that moves for reasons you can investigate |
What the table argues is not that sprawl is wrong. It is that sprawl is a risk management choice that has been sold as a deliverability choice, and those are different things. Spreading volume across forty domains genuinely limits blast radius. It does nothing whatsoever to make your mail more welcome, and it costs you the ability to detect the problem that causes domains to burn in the first place. Teams adopted it for the first property and assumed they were getting the second one free.
There is a middle position most estates have never tested, and it is where we push B2B SaaS teams with mature outbound programmes: run a small instrumented core alongside the sharded pool. Two or three domains carrying enough daily volume to produce real reputation data, sending the same sequences to the same segments as everything else, treated as the canary for the estate. When the core's spam rate moves, you have learned something about the copy and the list that the other thirty domains are experiencing silently. The core costs you concentration risk on a minority of your volume and buys you the only early warning available at this scale. For teams running our cold email engagements, that canary is now part of the standard build rather than an optimisation we suggest later.
Instrument The Estate Before A Receiver Does
None of this requires rebuilding the estate. Most of it is a week of provisioning and a change to how the weekly number gets produced.
Two things worth separating before anyone takes this to a planning meeting. The first is that the authentication layer is a solved problem and should be treated as one. Eighty four percent of From address domains carrying no DMARC record at all is a damning industry statistic, but it should not describe your estate, and if it does, fix that before worrying about anything in this argument. The record level work after the specification changed this year is a checklist, not a strategy, and it is the entry fee for the conversation rather than the conversation itself.
The second is that none of this makes bad outbound work. An estate with perfect instrumentation and a list nobody wanted to be on will still generate complaints, and the measurement will simply tell you so faster and with more precision. That is the actual value on offer here. Not fewer complaints, but knowing about them while you can still do something, rather than inferring them from a reply rate that fell off in the second week of the quarter.
Do this next. Open your sending platform, pull the daily volume for every domain in the estate, and add one column: one divided by that volume, as a percentage. Sort descending. Every domain where that number exceeds 0.30% is a domain whose complaint rate cannot be interpreted, and on most estates that will be every single row. Then count how many domains sit above roughly 100 Gmail bound messages a day. If the answer is none, you are not running a low complaint programme. You are running an unmeasured one, and the difference only becomes visible on the day placement drops and there is nothing in the history to explain it.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.