Ask a cold email team why their sequence runs seven steps and you get an answer about replies. Steps five through seven add a couple of points of reply rate, the reasoning goes, and a couple of points across ten thousand sends is real pipeline. The math is correct. It is also answering the wrong question, because reply rate is not the constraint that ends outbound programs. Complaint rate is, and no sequencing tool on the market shows it to you per step.
Two published numbers, from two sources with no relationship to each other, define the whole problem. Instantly's 2026 benchmark report, built on aggregated sends across thousands of workspaces between January 1 and December 18, 2025, found that 58 percent of all cold email replies arrive on step one. Google's own sender guidelines say to keep spam rate below 0.1 percent, and that at 0.3 percent or higher a sender becomes ineligible for mitigation support. Google also states that spam rate is calculated daily.
Put those together and a follow-up stops looking like free upside. It looks like a purchase, made with a budget that resets every twenty four hours, spent on the specific slice of your list that has already told you no by saying nothing.
The reply curve and the complaint curve point in opposite directions
Reply rate per step decays. That is the least controversial claim in outbound and every practitioner has watched it happen. The first email reaches a cold contact at their most neutral, and the people for whom the offer is obviously relevant say so immediately. Instantly's figure of 58 percent on step one is the cleanest public statement of it, and the corollary is that 42 percent of replies come from steps two and beyond combined.
Complaint rate per step does not decay. If anything it climbs, because irritation compounds where interest does not. The person who ignored your first three emails is measurably more likely to hit the spam button on the fourth than a fresh contact was on the first, and every additional touch keeps the same recipient in the denominator while adding fresh chances to be reported.
Reply share by step. Step one is the reported figure from Instantly's 2026 benchmark. Steps two and beyond distribute the reported 42 percent remainder as an illustrative geometric decay, not measured data, to show the shape of the curve rather than exact values.
Substitute your own decay curve, because the shape is the argument and not the figures. Whatever numbers you plug in, the back half of a seven step sequence is a large volume of sends chasing a small share of replies. The sends are the thing Google counts. The replies are the thing your dashboard counts. Those are two different ledgers and only one of them can shut you down.
Cold email sequence length is priced in complaints, not replies
Here is the conversion almost nobody performs. Google's thresholds are percentages, and percentages hide how few people are actually involved. Run the arithmetic on Google's published figures at volumes real outbound teams operate at and the margin becomes uncomfortably legible.
| DELIVERED TO GMAIL PER DAY | COMPLAINTS AT THE 0.1% GUIDANCE | COMPLAINTS AT THE 0.3% CEILING | WHAT THAT MEANS IN PRACTICE |
|---|---|---|---|
| 500 | 0.5 | 1.5 | A single complaint can put a light sending day over the recommended figure outright |
| 1,000 | 1 | 3 | Three irritated strangers in one day is the entire margin, and inboxes are not evenly distributed across days |
| 5,000 | 5 | 15 | Fifteen complaints across a full day of sending, which one badly matched segment can produce on its own |
| 10,000 | 10 | 30 | Enough headroom to hide a problem for weeks, which is why larger programs discover this late |
That table is division, not research. It is Google's own published percentages applied to four sending volumes, and it is worth doing on paper because percentages this small do not translate intuitively. Most teams reason about spam rate as a slow moving reputation score. Google describes it as a daily calculation, which makes a single bad afternoon a measurable event rather than something absorbed into a monthly average.
Step five is aimed at your worst matched prospects by design
This is the part that gets missed, and it is structural rather than tactical. Follow-ups do not go to a random sample of your list. They go, by definition, to everyone who did not reply. That population is not neutral. It has been filtered, by your own sequence, to remove the people who found the message relevant enough to answer.
So the deeper you go into a sequence, the worse the average fit of the person receiving it. Step one goes to your whole list, good matches and bad. Step five goes almost exclusively to the bad matches, plus a thin residue of good matches who were busy. You are spending your scarcest resource, the daily complaint allowance, on the segment with the highest complaint propensity and the lowest reply propensity, and you are doing it more aggressively with every additional step.
“A follow-up is not a second chance at the same audience. It is a first email to a worse one, assembled by your own sequence out of everyone who already passed.”
What Google measures, and what your sequencer shows you
The gap between those two views is the whole operational problem. Your sequencing platform reports opens, clicks, replies and bounces, attributed by step. Google reports authentication results, delivery errors and a spam rate, attributed by sending domain and calculated daily. Nothing in either system tells you which step generated which complaint, and that missing join is why sequence length keeps getting set on reply data alone.
There is a second filter stacked on top of all this now, which is that the inbox itself has started summarising and triaging before a human sees anything. We covered how AI inbox features act as a second gate on cold email, and the interaction with sequence length is not friendly. A message that gets bundled or deprioritised on step four still consumed a send, still sits in your daily volume, and still carries its full complaint risk if the recipient ever opens the bundle.
How to set cold email sequence length for the list you actually have
The honest answer is that there is no universal number, and the vendors quoting a four to seven step range are describing where reply rate flattens rather than where risk starts. Both numbers matter. Only one of them is in the benchmark.
Start from your allowance rather than from a best practice. Take your daily delivered volume to Gmail addresses, multiply by 0.001, and write down the whole number. That is how many people can report you on an average day before you are outside Google's recommended figure. Then look at your current sequence and ask which steps you would still run if each one had to justify itself against that number. For most teams sending to loosely qualified lists, the answer removes two steps and adds a week of spacing, and the reply loss is smaller than the reply chart implies because the deepest steps were carrying the least reply share to begin with.
Then fix the input rather than the sequence. A shorter sequence to a better list beats a longer sequence to a worse one on both ledgers at once, which is the only move in outbound that improves reply rate and complaint rate together. Everything else is a trade. This is why our cold email engagements spend more time on segmentation than on copy, and why B2B programs with narrow, well defined ICPs can run cadences that would destroy a broader sender.
The credibility side compounds too. Sender reputation is increasingly legible outside the inbox, and the same signals that make a domain trustworthy to a filter make a company trustworthy to a buyer checking you before they reply, which we walked through in how sender credibility now intersects with AI search.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.