Everybody I talk to about AI SDR performance wants to talk about the copy. Is the personalization good enough. Does it sound human. Should we run it through another model to take the edges off. Fair questions. Wrong questions, mostly.
A comparison published this year put 100,000 cold emails side by side, 50,000 written by AI and 50,000 written by people, across a window running from October 2025 into April 2026. The analysis of AI SDR versus human-written performance reports the AI side replying at 4.1 percent against 5.2 percent for humans. That is a gap. It is not a catastrophe, and if you stopped reading there you would conclude the machines are close and getting closer.
Keep reading and it stops being a copy story.
What the 100,000 email comparison actually found
Start with the bounce rate, because it is the least interesting number and the most important one. Six percent on both sides. Identical.
That means the two groups were working from list data of the same quality. Nobody handed the humans a clean list and the machines a scraped one. Whatever separates the two groups downstream, it is not data hygiene, and that single control is what makes the rest of the numbers worth arguing about.
| METRIC | AI-WRITTEN | HUMAN-WRITTEN | RELATIVE GAP |
|---|---|---|---|
| Bounce rate | 6% | 6% | None. List quality matched |
| Reply rate | 4.1% | 5.2% | Humans 27% higher |
| Positive replies | 1.4% | 2.1% | Humans 50% higher |
| Meetings booked | 0.7% | 1.1% | Humans 57% higher |
| Spam flag rate | 8% | 3% | AI 2.7x higher |
Look at the third column going down. Twenty seven percent. Fifty percent. Fifty seven percent. The gap widens at every stage of the funnel, and the relative-gap column is our arithmetic on the reported figures, not a separate claim from the study.
That shape matters. If AI copy were simply a bit blander than human copy, you would expect a roughly constant penalty all the way down. A slightly worse email gets slightly fewer replies, and those replies convert at about the same rate. That is not what happened. The AI-generated emails did not just get fewer replies, they got worse replies, and the replies they did get turned into meetings less often.
Something is filtering for quality of intent, not quality of prose. And the spam number tells you what.
AI SDR performance is losing on trust signals, not sentences
Eight percent of the AI-written emails got flagged as spam. Three percent of the human ones did. Same lists, same bounce rate, nearly three times the flag rate.
Spam filters are not literary critics. They are not sitting there deciding your second paragraph reads like it was generated. They are looking at sending patterns, domain reputation, volume ramps, engagement history, and how closely a message resembles other messages hitting the same infrastructure. That is a behavioural fingerprint, and AI-assisted outbound has a very distinctive one, because the whole point of automating it is to do more of it, faster, from more addresses.
So the copy is a red herring. The machine is not being punished for how it writes. It is being punished for how it operates: higher volume, tighter intervals, more templated structure across a shared sending footprint, and a reply signal too thin to build a positive reputation on. Once you land in a filter, the recipients you needed most never see the message, and those are disproportionately the ones at larger companies with the strictest mail policies. Which is exactly where your meetings were going to come from.
None of this says do not use AI in outbound. It says the thing AI is good at, which is producing volume, is the thing that is most expensive in this channel right now. That is a genuinely awkward finding for the category and it deserves to be said plainly rather than buried. We made a version of this argument about what sender credibility now costs you, and this dataset is the operational evidence for it.
Cadence beat every copy variable in the test
Here is the part that should have been the headline.
Moving the send interval from one day to three days took inbox placement from 71 percent to 93 percent. Twenty two percentage points. No new copy, no new personalization layer, no model upgrade. Just waiting longer between touches.
Inbox placement by send interval, as reported in the 100,000 email analysis
Sit with the scale of that for a second. Twenty two points of inbox placement is the difference between a sequence that works and a sequence that quietly does not, and it dwarfs any plausible lift from rewriting a subject line. The study's own framing is that cadence beats content, and on these numbers that is hard to argue with.
It is also the least glamorous finding imaginable, which is probably why nobody is selling it. There is no product in wait longer. There is no demo. You cannot put a slider on it. So the entire conversation drifts back to copy, where the tooling is, and away from the setting that actually moved the number.
“The lever that moved most was the one nobody can sell you. That is usually how it goes.”
The mechanism is not mysterious. Tighter intervals concentrate volume, and concentrated volume from a young or lightly-warmed domain is one of the loudest signals a filter has. Spreading the same sequence over more days lowers the peak rate per mailbox per hour, gives positive engagement time to register between touches, and stops the pattern from looking like a burst. It is the same physics behind monitoring your sending reputation directly rather than inferring it instead of guessing from open rates that were never reliable anyway.
How to read a vendor-adjacent benchmark without getting played
Now the caveats, because I would rather you take this seriously than take it literally.
This is not a peer-reviewed study. It is an analysis compiled by an agency from aggregated data across several outbound platforms plus their own databases. That is a legitimate way to build a dataset and it is also a way to build one with its thumb on the scale, because the people who assemble these numbers usually have a service to sell on the back of them. Take the directions seriously and hold the decimal places loosely.
Three specific things I would want before treating any of these figures as settled. First, how were the AI and human groups assigned, because if the human-written emails went to warmer or better-researched accounts then the whole comparison collapses. Second, what counts as a positive reply, since that definition swings by a factor of two between vendors. Third, whether the cadence finding was tested inside both groups or only observed across them, because a correlation between slower sending and better placement could partly be a proxy for more careful operators.
That is not scepticism for its own sake. It is that outbound benchmarks get quoted for years after the conditions that produced them have changed, and the ones that spread fastest are the ones with the most quotable single number. The 4.1 against 5.2 will travel. The 71 to 93 will not, and it is the one worth having.
The AI SDR performance fixes worth doing this month
Four things, roughly in order of how much they will move.
Change your send interval first. Go from one day to three between touches on one live sequence, hold everything else constant, and measure inbox placement with seed accounts rather than open tracking. This is a settings change. It costs you nothing but calendar time, and on the evidence above it is the highest-leverage thing available to you this month.
Then split your reporting so the meeting rate leads and the reply rate follows. If your outbound dashboard opens on reply rate, the widening funnel gap this data describes is invisible to you, and you will keep making decisions on the number that flatters the program. The gap between a 27 percent difference and a 57 percent difference is exactly the gap between renew this and kill this.
Third, audit where AI is actually sitting in your process. Using it to research an account, summarize a trigger, or draft a first pass a human then rewrites is a different thing from using it to generate and send at volume. The first shows up as leverage. The second shows up as an eight percent spam flag rate. Most teams describe themselves as doing the first and, when you look at the sending logs, are doing the second.
Fourth, put your domain and mailbox architecture on the same review cadence as your copy. Sending domains, per-mailbox volume, warmup state, and reputation monitoring get reviewed monthly in the cold email programs we run, because these are the variables that decide whether any of the copy work gets seen at all. For tech and SaaS teams selling into enterprises with strict mail policy, this is not a marginal concern, it is most of the game.
You are not behind because your AI copy is not good enough yet. That framing has cost a lot of teams a lot of quarters. The copy is fine. The volume it enables, sent on the infrastructure most people are sending it on, at the intervals most people are using, is what is costing you meetings. Fix the operating envelope first. Then, if you still want to argue about the second paragraph, at least somebody will be reading it.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.