Something Inc.LoginSchedule a free consultation
STRATEGY

Half your deliverability checklist has no number behind it

Some cold email deliverability rules come with a published threshold from the mailbox provider. Most come from a vendor blog. Only one of those groups is worth arguing about.

TTTyler TruffiManaging Partner · AUG 17, 2026 · 10 MIN READ

Ask five cold email vendors what your daily sending limit per inbox should be and you will get five confident answers, none of which cite anything. Ask them what Gmail's spam complaint threshold is and, if they know their job, you will get one answer, because Google published it. That difference is the most useful sorting tool in outbound, and almost nobody applies it.

This is not an argument that unpublished rules are wrong. Plenty of them encode real operational experience. It is an argument that you should know which of your cold email deliverability rules you can point at a primary source for, because that determines how hard you should fight for them when someone senior wants to override you.

Two piles: published numbers and everything else

Pile one: a mailbox provider stated a specific requirement or threshold in its own documentation, with a compliance date. These are enforceable, testable, and non-negotiable. When you fail them the failure is legible, often with a specific rejection code attached.

Pile two: everything else. Warmup schedules, per-inbox daily caps, domain-to-inbox ratios, optimal send windows, the exact number of days before a follow-up. Some of this is well-grounded pattern recognition from people who send a lot of mail. Some of it is a number someone picked in 2021 that got copied into forty blog posts. From the outside the two look identical, because both arrive as a confident integer.

THE SORTING RULEIf you cannot name the provider document and the date it was published, the rule belongs in pile two. That does not mean stop following it. It means stop treating disagreement about it as heresy, and start testing it on your own data.

What Google and Microsoft actually published

The bulk sender requirements Google and Yahoo introduced set a clear bar, and it is genuinely short. Senders above 5,000 messages a day to Gmail addresses must authenticate with SPF, DKIM, and DMARC, with DMARC at a minimum policy of p=none. Marketing and subscription mail needs one-click unsubscribe. And the spam rate bar is explicit: stay below 0.10% for good standing, with 0.30% named as the level that triggers rejections. Enforcement rolled out through 2024, starting with warnings and increased filtering in February and moving to outright rejection of non-compliant bulk mail through the spring.

0.10%
Gmail spam complaint rate for good standing (published threshold)
0.30%
spam rate level Google names as triggering rejections
5,000/day
bulk sender threshold at both Google and Microsoft

Microsoft's version, which we covered when it landed, followed the same shape: announced April 2, 2025, enforced from May 5, 2025, applying to senders above 5,000 messages a day to Outlook.com, Hotmail, and Live addresses, requiring SPF, DKIM, and DMARC at p=none with alignment, and returning a specific rejection code, 550 5.7.515, when you fail. Since then Microsoft has expanded ARC handling so authentication survives relays, and DKIM has moved from best practice to effectively required for reliable delivery.

Notice what is not in that list. No provider published a per-inbox daily cap. None published a warmup duration. None published an acceptable number of sending domains, a ratio of inboxes per domain, or a recommended ramp curve. The published rules are about authentication, complaint rate, and unsubscribe mechanics. Everything operational that outbound teams argue about most is absent from the primary sources entirely.

RULEPUBLISHED BY A PROVIDER?HOW TO TREAT IT
SPF, DKIM, DMARC at p=none above 5,000/dayYes, Google and Microsoft, with datesNon-negotiable. Fix before anything else.
Spam complaint rate under 0.10%Yes, Google, explicit figureYour primary health metric. Monitor continuously.
One-click unsubscribe on bulk mailYes, Google and YahooNon-negotiable for anything list-based.
Rejection code 550 5.7.515 on failureYes, Microsoft, May 2025Diagnostic. Tells you exactly what broke.
Daily sends per inbox (20, 30, 50...)NoTest on your own data. Vendor defaults are guesses.
Warmup duration before real sendsNoDirectionally sound, specific numbers unsourced.
Inboxes per sending domainNoOperational preference dressed as a rule.
Optimal send day and hourNoTest it. Most of the published claims are aggregate B2C data.

There is one more thing worth pulling out of the published set, because teams consistently misread it. The 5,000 per day threshold is a bulk sender definition, not a safety limit. It marks the point above which the authentication and unsubscribe requirements become mandatory. It does not imply that 4,999 messages a day is safe, or that staying under it exempts you from filtering. Plenty of senders well below the bulk threshold land in spam constantly, because complaint rate and authentication quality apply to everyone and the threshold only governs which requirements are formally enforced. Reading it as a speed limit is one of the most common and most expensive misreadings in outbound, and it leads teams to split volume across more domains for no defensible reason. Compare the sender requirements Google documents against whatever your platform's onboarding told you, and the gap is usually instructive.

The rules with no number behind them

The pile two rules deserve a fair hearing, because dismissing them wholesale is its own kind of sloppy. A team sending from a brand new domain at 200 messages a day on day one will have a bad time, and no provider needed to publish a document for that to be true. Gradual ramp is real. Volume sensitivity is real. The problem is precision without provenance: not the idea of a ramp, but the confident claim that it must be exactly fourteen days, or that 30 per inbox is safe while 35 is reckless.

1Unsourced numbers get copied, not testedA specific integer in a vendor blog gets quoted by the next twenty blogs. Nothing about that chain constitutes evidence, but the repetition reads as consensus.
2Vendor defaults encode vendor economicsA platform that sells inboxes has a structural reason to recommend more inboxes at lower volume each. That may still be correct advice. It is not disinterested advice.
3Your domain history is not the averageAggregate guidance cannot know whether your domain has ten years of clean sending or was registered last month. The right cap for you is a function of history the rule cannot see.
4The published rules are the ones with error codesWhen you break a pile one rule, you get a rejection code naming the failure. When you break a pile two rule, you get worse placement and no explanation. That asymmetry is why teams over-index on folklore.

That last point is the honest reason folklore thrives here. Deliverability failure is mostly silent. Mail lands in spam without telling you, reply rates sag, and there is no log line that says the cause. In an environment with that little feedback, a confident rule is emotionally useful even when it is unfounded. It gives a team something to comply with. The cost is that nobody ever tests it, and a decade of untested caution compounds into a program sending a fraction of what it safely could.

Why cold email deliverability rules get invented

There is a structural reason the unpublished pile keeps growing. Providers publish the minimum enforceable bar because that is what they can commit to. The actual filtering decision is made by a machine learning system that updates continuously and that no provider will ever document, because documenting it is a specification for gaming it. So the gap between the published floor and the real behavior is permanent, and the vendor ecosystem fills it with numbers.

Which means the two piles are not evidence versus nonsense. They are what the provider will commit to in writing versus what practitioners inferred from watching outcomes. Inference is legitimate. Inference presented as a published standard is not. The distinction sounds pedantic until the quarter your outbound program gets restructured around a number nobody can source, and the reply rate distribution tells you the change did nothing.

Providers publish the floor, not the algorithm. Everything between the floor and the filter is someone's inference, and inference should be tested, not obeyed.

This also explains why deliverability advice ages so badly. Pile one rules change on announced dates with compliance windows, the way Microsoft's did between April and May 2025. Pile two rules change silently, if at all, because there was never a mechanism to update them. A warmup schedule written for 2022 filtering is still circulating unchanged, and it looks exactly as authoritative as it did then.

How to run the sort on your own checklist

This is an afternoon of work and it usually surprises people. Write out every deliverability rule your team currently follows, then put a source column next to it and try to fill it in with a provider document and a date. Most teams get through the authentication rules quickly, then hit a wall around the operational caps, and discover that the numbers governing the largest part of their sending capacity trace back to a platform onboarding screen nobody has questioned since setup. That is not a scandal. It is just an unexamined default, and unexamined defaults are the cheapest thing in any program to go fix.

One caution before you start relaxing caps. Sorting rules by evidence is not permission to ignore the unsourced pile, and a team that responds to this piece by doubling every send limit at once has misread it badly. The correct output is a testing queue ordered by cost, not a bonfire. Pile two rules are hypotheses, and hypotheses get tested one at a time against a metric you are already watching, so that when something moves you know what moved it.

COMPLY
Audit compliance against pile one firstSPF, DKIM, DMARC alignment, one-click unsubscribe, and current spam rate. If any of these are wrong, nothing in pile two matters until they are fixed.
MONITOR
Instrument complaint rate continuously0.10% is the only hard number you get. Treat it as a live dashboard metric, not a quarterly check, and know your blind spots where provider feedback loops go dark.
TEST
Pick one pile two rule and test itChoose the cap costing you the most volume. Run a controlled increase on a subset of inboxes for four weeks, watching complaint rate and reply rate. One test beats ten opinions.
DOCUMENT
Write the source column into your runbookEvery rule gets a citation or gets flagged as inferred. New team members then inherit the distinction instead of inheriting the folklore.

The monitoring point deserves emphasis, because complaint rate is the one published metric that can go dark on you without warning. When provider feedback data stops flowing, as we saw during the SNDS blackout, teams that had built their entire health picture on one dashboard discovered they had no independent read on their own sending. Redundancy in measurement is not paranoia when the single available number is also the only enforceable one.

For teams running outbound at real scale, the payoff from this sort is usually volume. Most programs are throttled well below what their domain history and complaint rate would support, because the caps were set by a vendor default years ago and never revisited. In our FMS Investor engagement the gains came from getting the fundamentals right and then measuring, not from adopting a stricter version of everyone else's rules. Cautious defaults are cheap to adopt and expensive to keep.

Do this next: pull up your current spam complaint rate and put it against 0.10%. If you cannot find that number in under five minutes, that is the actual finding, and it outranks every other item on your list. Fix the measurement, confirm authentication is clean end to end, and only then start testing the rules nobody can source. The full sequencing for that lives in our cold email deliverability playbook, and it is the same sequence we run at the start of every cold email engagement: published rules first, inferred rules second, opinions never.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

TT
Tyler TruffiMANAGING PARTNER, SOMETHING INC.

Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.