Something Inc.LoginSchedule a free consultation
STRATEGY

Cold email sequence length is a complaint rate decision

Every sequencing tool prices follow-ups in replies. Gmail prices them in complaints, daily, against a threshold most teams have never converted into an actual number of annoyed strangers.

STRATEGYCOLD EMAILDECISION PIECE

Ask a cold email team why their sequence runs seven steps and you get an answer about replies. Steps five through seven add a couple of points of reply rate, the reasoning goes, and a couple of points across ten thousand sends is real pipeline. The math is correct. It is also answering the wrong question, because reply rate is not the constraint that ends outbound programs. Complaint rate is, and no sequencing tool on the market shows it to you per step.

Two published numbers, from two sources with no relationship to each other, define the whole problem. Instantly's 2026 benchmark report, built on aggregated sends across thousands of workspaces between January 1 and December 18, 2025, found that 58 percent of all cold email replies arrive on step one. Google's own sender guidelines say to keep spam rate below 0.1 percent, and that at 0.3 percent or higher a sender becomes ineligible for mitigation support. Google also states that spam rate is calculated daily.

58%
of cold email replies land on step one, per Instantly's 2026 benchmark report
0.1%
spam rate Google's sender guidelines tell you to stay below
0.3%
spam rate at which Google says a sender is ineligible for mitigation support
Daily
frequency at which Google states spam rate is calculated, not weekly or monthly

Put those together and a follow-up stops looking like free upside. It looks like a purchase, made with a budget that resets every twenty four hours, spent on the specific slice of your list that has already told you no by saying nothing.

The reply curve and the complaint curve point in opposite directions

Reply rate per step decays. That is the least controversial claim in outbound and every practitioner has watched it happen. The first email reaches a cold contact at their most neutral, and the people for whom the offer is obviously relevant say so immediately. Instantly's figure of 58 percent on step one is the cleanest public statement of it, and the corollary is that 42 percent of replies come from steps two and beyond combined.

Complaint rate per step does not decay. If anything it climbs, because irritation compounds where interest does not. The person who ignored your first three emails is measurably more likely to hit the spam button on the fourth than a fresh contact was on the first, and every additional touch keeps the same recipient in the denominator while adding fresh chances to be reported.

Step 1, reported figure58%
Step 2, illustrative17%
Step 3, illustrative11%
Step 4, illustrative7%
Steps 5 to 7 combined, illustrative7%

Reply share by step. Step one is the reported figure from Instantly's 2026 benchmark. Steps two and beyond distribute the reported 42 percent remainder as an illustrative geometric decay, not measured data, to show the shape of the curve rather than exact values.

Substitute your own decay curve, because the shape is the argument and not the figures. Whatever numbers you plug in, the back half of a seven step sequence is a large volume of sends chasing a small share of replies. The sends are the thing Google counts. The replies are the thing your dashboard counts. Those are two different ledgers and only one of them can shut you down.

Cold email sequence length is priced in complaints, not replies

Here is the conversion almost nobody performs. Google's thresholds are percentages, and percentages hide how few people are actually involved. Run the arithmetic on Google's published figures at volumes real outbound teams operate at and the margin becomes uncomfortably legible.

DELIVERED TO GMAIL PER DAYCOMPLAINTS AT THE 0.1% GUIDANCECOMPLAINTS AT THE 0.3% CEILINGWHAT THAT MEANS IN PRACTICE
5000.51.5A single complaint can put a light sending day over the recommended figure outright
1,00013Three irritated strangers in one day is the entire margin, and inboxes are not evenly distributed across days
5,000515Fifteen complaints across a full day of sending, which one badly matched segment can produce on its own
10,0001030Enough headroom to hide a problem for weeks, which is why larger programs discover this late

That table is division, not research. It is Google's own published percentages applied to four sending volumes, and it is worth doing on paper because percentages this small do not translate intuitively. Most teams reason about spam rate as a slow moving reputation score. Google describes it as a daily calculation, which makes a single bad afternoon a measurable event rather than something absorbed into a monthly average.

THE REFRAMEYou do not have a spam rate. You have a daily complaint allowance, denominated in individual human beings, and every step you add to a sequence spends more of it on the contacts least likely to convert. Setting cold email sequence length without knowing that number is setting it blind, and it is the reason a program can look healthy for a quarter and then fall off a cliff in a week.

Step five is aimed at your worst matched prospects by design

This is the part that gets missed, and it is structural rather than tactical. Follow-ups do not go to a random sample of your list. They go, by definition, to everyone who did not reply. That population is not neutral. It has been filtered, by your own sequence, to remove the people who found the message relevant enough to answer.

So the deeper you go into a sequence, the worse the average fit of the person receiving it. Step one goes to your whole list, good matches and bad. Step five goes almost exclusively to the bad matches, plus a thin residue of good matches who were busy. You are spending your scarcest resource, the daily complaint allowance, on the segment with the highest complaint propensity and the lowest reply propensity, and you are doing it more aggressively with every additional step.

A follow-up is not a second chance at the same audience. It is a first email to a worse one, assembled by your own sequence out of everyone who already passed.
The non-responder pool is adversely selectedEvery step removes the people who liked the message and keeps the people who did not. By step four the recipient list is a concentrate of poor fit, and its behaviour should be modelled that way rather than as a slightly less responsive version of your original list.
Complaint propensity is not evenly spreadA small number of contacts account for most spam reports, and they are disproportionately in the non-responder pool. Which means the marginal step does not add a little complaint risk to everyone. It adds a lot to a few, and those few are enough to move a daily percentage.
List quality changes the answer entirelyA tightly qualified list of two hundred accounts can carry seven steps without trouble, because the non-responder pool is still full of genuinely good fits. A loosely filtered list of twenty thousand cannot, and this is where vetting your prospecting data vendor turns from a procurement question into a deliverability one.
Cadence spacing is a partial mitigation, not a fixSpreading seven steps over eight weeks lowers the daily denominator pressure and gives irritation time to fade. It does not change who the later steps are aimed at. We have argued before that cadence discipline beats copy quality in automated outbound, and this is the sharpest version of why.

What Google measures, and what your sequencer shows you

The gap between those two views is the whole operational problem. Your sequencing platform reports opens, clicks, replies and bounces, attributed by step. Google reports authentication results, delivery errors and a spam rate, attributed by sending domain and calculated daily. Nothing in either system tells you which step generated which complaint, and that missing join is why sequence length keeps getting set on reply data alone.

01Split your sending domains by sequence depth, not just by volumeMost teams shard domains to spread volume. Shard them to isolate risk instead: steps one and two on one set, steps three and beyond on another. You lose nothing operationally and you gain the only per-step complaint signal available to you, because each domain now reports a spam rate that maps to a known part of the sequence.
02Treat your bulk sender status as a threshold you can cross accidentallyGoogle defines a bulk sender as anyone sending close to 5,000 messages or more to personal Gmail accounts in a 24 hour period. A seven step sequence multiplies your effective daily volume against the same contact base, and teams cross that line through sequence depth without ever increasing their list size.
03Assume the unsubscribe rules apply to youGoogle's guidelines attach one-click unsubscribe and a 48 hour processing window to marketing and promotional messages. A templated sequence sent to a purchased list is promotional, whatever your sales team calls it. Arguing otherwise is a legal position, not a deliverability one, and the engine filtering you does not read your argument. The country by country compliance picture makes the same point from the regulatory side.
04Read complaint rate against delivered volume, not sent volumeBounces leave your sent count but not your problem. A list with a high bounce rate inflates the sends you think you made and shrinks the denominator Google is dividing by, so the same absolute number of complaints reads worse. Clean the list before you lengthen the sequence, in that order.

There is a second filter stacked on top of all this now, which is that the inbox itself has started summarising and triaging before a human sees anything. We covered how AI inbox features act as a second gate on cold email, and the interaction with sequence length is not friendly. A message that gets bundled or deprioritised on step four still consumed a send, still sits in your daily volume, and still carries its full complaint risk if the recipient ever opens the bundle.

How to set cold email sequence length for the list you actually have

The honest answer is that there is no universal number, and the vendors quoting a four to seven step range are describing where reply rate flattens rather than where risk starts. Both numbers matter. Only one of them is in the benchmark.

Start from your allowance rather than from a best practice. Take your daily delivered volume to Gmail addresses, multiply by 0.001, and write down the whole number. That is how many people can report you on an average day before you are outside Google's recommended figure. Then look at your current sequence and ask which steps you would still run if each one had to justify itself against that number. For most teams sending to loosely qualified lists, the answer removes two steps and adds a week of spacing, and the reply loss is smaller than the reply chart implies because the deepest steps were carrying the least reply share to begin with.

Then fix the input rather than the sequence. A shorter sequence to a better list beats a longer sequence to a worse one on both ledgers at once, which is the only move in outbound that improves reply rate and complaint rate together. Everything else is a trade. This is why our cold email engagements spend more time on segmentation than on copy, and why B2B programs with narrow, well defined ICPs can run cadences that would destroy a broader sender.

The credibility side compounds too. Sender reputation is increasingly legible outside the inbox, and the same signals that make a domain trustworthy to a filter make a company trustworthy to a buyer checking you before they reply, which we walked through in how sender credibility now intersects with AI search.

DO THIS NEXTBefore you touch copy this week, do the two minute version. Pull your delivered volume to Gmail addresses for the last thirty days, find your worst single day, multiply it by 0.001 and by 0.003, and put those two whole numbers at the top of your outbound dashboard where the reply rate currently sits. If the first number is under five, your sequence length is already a risk decision whether or not anyone on the team has treated it as one. Google's published sender guidelines are the source for both thresholds, and they take ten minutes to read against your own setup.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

TT
Tyler TruffiMANAGING PARTNER, SOMETHING INC.

Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.