Somebody on your team is about to put a cold email reply rate benchmark in a slide. They found it in a blog post, it has a decimal point in it, and it is going to sit next to your actual number in a red or green box. I want to talk you out of that box before it costs someone their quarter.
Here is the situation. Three email platforms published response-rate studies covering roughly the same period, all with real sample sizes, all reputable, none of them lying. Belkins looked at 7,530,489 cold emails sent across 2025 and reported an average reply rate of 0.45%. Instantly, in its 2026 benchmark report published January 12, put the average at 3.43% across what it describes as billions of interactions from Jan 1 to Dec 18, 2025. Woodpecker, drawing on 20 million-plus sales emails from more than a thousand customers in 52 countries, reported campaigns with three to five follow-up steps replying at 8.3%.
Eighteen times, low to high. Same year. Same channel. Same word on the label. If you picked the wrong one, your perfectly healthy campaign looks like a catastrophe, or your genuinely broken campaign looks like a triumph. Neither of those is a small mistake.
Three studies, three different worlds
Put them side by side and the gap stops looking like a mystery. It starts looking like three teams answering three different questions and giving the answers the same name.
| STUDY | SAMPLE | PERIOD | REPORTED REPLY RATE | WHAT IT PLAUSIBLY COUNTS |
|---|---|---|---|---|
| Belkins (Margaret Lee, Jun 26, 2026) | 7,530,489 emails, 34,393 replies | Jan-Dec 2025 | 0.45% average | Tracked replies against total sends, agency-run campaigns at volume |
| Instantly 2026 benchmark report (Jan 12, 2026) | Billions of interactions, thousands of workspaces, anonymized | Jan 1 - Dec 18, 2025 | 3.43% average, 5.5%+ top quartile, 10.7%+ elite | Platform-side replies across a self-selected active user base |
| Woodpecker (Margaret Sikora, updated Jun 23, 2026) | 20M+ emails, 1,000+ customers, 52 countries | Not stated as a single window | 8.3% with 3-5 follow-up steps, 4.1% with none | Campaign-level reply rate, segmented by sequence depth |
Look at the denominators first. A reply rate calculated per email sent and a reply rate calculated per contact reached are not the same number, and a five-step sequence makes that difference roughly five-fold on its own. Then look at the numerator. All replies, positive replies, tracked replies, and replies-that-were-not-an-auto-responder are four populations, and the studies do not agree on which one they are reporting.
Why the cold email reply rate benchmark moves 18x
There is a second layer under the definitions, and it is the more interesting one. Belkins is an agency. Its dataset is what happens when a services business runs high-volume outbound on behalf of clients who bought a meetings number. Instantly is a platform. Its dataset is what happens across everyone who logs in, which skews toward people actively tuning campaigns because they are paying for a tool to do it. Woodpecker's segmented figures come from campaigns, not sends, and campaigns are a unit that operators shape deliberately.
None of those populations is wrong. All of them are unrepresentative of you specifically, in different directions. And the direction matters: the agency number is a floor built from scale, the platform number is a middle built from engaged users, the campaign number is a ceiling built from the sequences people bothered to structure well.
Belkins reply rate by recipient seniority and company size, from its 7.5M-email dataset (2025)
Those bars are hundredths of a percent scaled for readability — 0.57% for founders, 0.22% for the largest enterprises. Read them as ratios, not absolute heights. Inside one consistent dataset, seniority moves reply rate by 1.8x and company size moves it by 3.3x. That is the whole argument in one chart. If a single variable inside a single study swings the number by more than three times, a cross-study comparison against a different variable set is not a measurement. It is a coin flip with a decimal point.
The geography spread in the same dataset makes the point harder still: Poland at 1.43%, Ireland at 0.74%, the United States at 0.51%. Food and beverage as a target industry came in at 3.47%, roughly eight times the overall average. Every one of those slices would be a defensible headline if someone wanted a number to sell you something.
Selection bias is doing most of the work
Notice who publishes these studies. Platforms and agencies publish benchmarks because benchmarks generate links and demos. That is not a scandal — we publish data too, and we would rather the industry had more of it, not less. But it does mean the sampling frame is always the vendor's own customers, and the vendor's own customers are people who already decided outbound was worth paying for.
The practical version of that: nobody in any of these datasets is the team that sent 200 emails from a brand-new domain with no authentication and gave up in week three. That team exists, it is common, and it is not in the sample. Which means every published benchmark is, structurally, an optimistic reading of a survivor population.
“A benchmark you did not build and cannot audit is a story about someone else's customers, told in your metric's name.”
We have made this argument before about the open rate, which stopped being interpretable the moment Apple's Mail Privacy Protection started pre-loading tracking pixels. The reply rate is in better shape because a reply requires a human decision, but only if you control what counts as one. Otherwise you have imported someone else's definition and lost the one property that made the metric worth having.
The two findings that replicate everywhere
Here is the useful part. When three datasets with wildly different absolute numbers agree on a direction, that direction is probably real. Two things clear that bar.
What the spread does to your pipeline forecast
This is not an academic complaint, because the reply rate is an input to a forecast and the forecast is an input to a headcount decision. Run the arithmetic and the harm becomes obvious. Take a program sending 20,000 emails a quarter, with a third of replies turning into a meeting and a fifth of meetings turning into an opportunity. At the 0.45% figure that is 90 replies, 30 meetings and six opportunities. At 3.43% it is 686 replies, 229 meetings and 46 opportunities. Same activity, same team, same quarter — a difference of forty opportunities, purely from which blog post the planning spreadsheet borrowed its assumption from.
Building a cold email reply rate benchmark you can trust
The version that works is unglamorous and takes one quarter. Define the metric once, in writing: replies from a human, per contact reached, excluding automatic responses, counted within fourteen days of the last touch. Write down whether a referral to a colleague counts. Write down whether a negative reply counts, because it should — a no is a human decision and it tells you the targeting worked even when the offer did not.
Then segment before you average. One blended number across seniority, company size, industry and geography hides exactly the variation the studies just showed you is enormous. Report reply rate by segment against that segment's own prior quarter, and the comparison becomes real, because the only variable that changed is what you did. This is the same discipline that makes departmental benchmarks readable rather than decorative, and it is the difference between a metric that survives a leadership review and one that gets argued with.
Give it a floor of statistical honesty too. A segment with 200 sends and three replies is not a 1.5% reply rate, it is noise with a percent sign. Set a minimum volume before a segment gets a number at all, and hold to it even when the empty cell is uncomfortable. Our cold email engagements run this way from week one, and our B2B work uses the same segment-first structure, because the alternative is quarterly arguments about whose blog post was right.
Do this next
Open whatever deck has the industry benchmark on it and delete the benchmark. In its place, put your own number from last quarter, segmented by seniority and company size, with the definition written underneath in one sentence. Then build the follow-up steps you are missing, because that is the one intervention all three studies endorse and the cheapest thing on this list.
You are not behind. You are probably measuring a different thing than the post you read, and that is fixable in an afternoon. For the adjacent arguments, we have written on why reply tracking beats open tracking, on how reply rates split by department, and on what the reply-rate data actually supports as best practice. Belkins published its full 2025 breakdown if you want the segment tables in the original.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.