Woodpecker's cold email dataset, drawn from more than 20 million sent emails and last updated June 23, 2026, has a finding most teams' actual practice contradicts directly: lists under 50 contacts reply at 5.8%, while lists over 1,000 contacts reply at 2.1%. That's not a small gap. It's nearly three times the reply rate, running in the opposite direction of what most outbound teams optimize for, which is almost always more contacts, not fewer.
The cold email list size number most teams get backwards
Ask most SDR leaders how to get more pipeline out of cold email and the instinct is almost always the same: buy or scrape a bigger list, add more contacts to the sequence, increase total volume. Woodpecker's data, one of the largest publicly available cold email datasets we've seen this year, argues the opposite instinct is the correct one. Reply rate doesn't hold steady as list size grows. It falls, and it falls by a lot.
Reply rate by list size (Woodpecker, 20M+ emails, updated Jun 23, 2026)
Read literally, a list under 50 contacts replies at 5.8%, nearly three times the 2.1% a list over 1,000 contacts pulls. That's not a rounding difference you can explain away with sample-size noise, on a dataset this large. It's a structural relationship between list size and reply quality that most teams' actual sending behavior runs directly against.
Personalization moves the same number, harder
The same dataset found an even bigger gap on a related lever: personalized sequences averaged a 17-18% reply rate, versus 7-9% for unpersonalized ones. That's roughly double, on top of everything list-size segmentation already buys you. The two findings aren't separate levers by accident. They're the same underlying mechanism showing up twice: the more precisely a message is built for the specific person receiving it, whether that precision comes from a tightly scoped list or from message-level personalization, the better it performs. A giant, loosely-targeted list is the least precise version of outbound you can run, and a personalized message sent to a giant, loosely-targeted list is fighting its own targeting the whole way through. It's the same instinct behind the finding that 58% of replies come from a sequence's first email: effort spent sharpening the highest-leverage moment beats effort spent adding more volume around it.
| LEVER | LOWER-PRECISION VERSION | HIGHER-PRECISION VERSION | REPLY RATE GAP |
|---|---|---|---|
| List size | 1,000+ contacts | Under 50 contacts | 2.1% → 5.8% |
| Message personalization | Unpersonalized sequence | Personalized sequence | 7-9% → 17-18% |
“A giant list and a personalized message are fighting each other, not working together. Precision is the variable that actually moves the number, whichever lever you pull it with.”
Why this happens: signal dilution, not laziness
It's tempting to read the list-size finding as 'small lists get more attention per contact, so of course they do better,' and stop there. That's part of it, but the more useful explanation is signal dilution. A tightly scoped list of under 50 contacts is almost always built around a specific, real signal: a recent funding round, a shared trigger event, a genuinely narrow ICP fit. A list of 1,000-plus is, by construction, harder to build around a single sharp signal, because holding that much precision at that scale usually isn't feasible with the research time available. The list gets bigger by relaxing the criteria, and reply rate falls because the message that fit a narrow signal now has to work for a much fuzzier one.
This connects directly to something we've written about before: catch-all domains and list decay quietly drag down list quality the same way oversized lists do, by diluting the share of contacts who are actually a fit. A 5,000-contact list assembled quickly, with weak validation and loose ICP criteria, is accumulating the same kind of dead weight from two directions at once: contacts who were never a great fit, and contacts whose data has simply gone stale since the list was built.
The open rate caveat, and the bounce rate that's easy to miss
Two more numbers from the same dataset are worth flagging before anyone builds a dashboard around them. Open rate across the study ranged from 27.7% to 44%, and Woodpecker itself flags that range as inflated by Apple Mail Privacy Protection, which pre-fetches images and can register a message as opened whether or not a human actually looked at it. That's a real, disclosed limitation on one of the most commonly tracked cold email metrics, and it's a good reminder that reply rate, not open rate, is the number worth building strategy around. A list can show a great open rate and still be a weak, low-signal list if replies don't follow.
Average bounce rate across the full dataset sat at 5.1%, well above the sub-2% ceiling most deliverability guidance treats as the safe operating range, the same range we've referenced when covering Gmail's tightened spam-complaint threshold. A 5.1% average bounce rate across 20 million emails means a meaningful share of the senders in this dataset are working from stale or poorly validated lists, which likely drags the large-list reply-rate figure down further still: a 1,000-plus contact list assembled without rigorous validation is accumulating both weak-fit contacts and dead addresses at the same time. Tightening a list for signal, the fix this piece has focused on, also tends to tighten it for deliverability, since the same research effort that finds a real trigger event usually also confirms the contact is current and reachable.
How to apply the cold email list size data without cutting volume
The naive read of this data is 'send to fewer people,' which isn't actually the right takeaway for a team with a real pipeline number to hit. The right takeaway is: stop treating list size as the lever that gets you there, and rebuild list construction around tighter segments, even if you're running more of them in parallel to hit the same total volume.
Do this next
Pull your current active lists and check their size distribution against reply rate. If you're already segmenting, you likely already see some version of this pattern in your own data; confirm it before you restructure anything. If you're running a small number of large, broad lists, pick your single biggest one this week and split it into at least three tighter segments, each built around a distinct, real signal, and run them in parallel rather than as one campaign.
This doesn't mean shrinking total outbound volume. It means building that volume out of more, smaller, sharper pieces instead of fewer, bigger, blunter ones. The 3.43% average reply rate most teams benchmark against assumes the current mix of list sizes across the industry. A team that shifts its list construction toward the smaller, higher-signal end of that distribution has a real, sourced reason to expect it can beat that average, not just hope to. If this changes how your team allocates research time, it's worth revisiting how cold email lead generation work gets staffed generally: list construction stops being a one-time setup task and becomes an ongoing, measurable part of the program, the same way we've argued reply rates themselves should be tracked by segment, not just as one blended monthly average.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.