Something Inc.LoginSchedule a free consultation
STRATEGY

Your AI SDR does not have a writing problem

A 100,000 email comparison of AI-written and human-written cold email found the AI side losing on replies, losing worse on meetings, and getting flagged as spam nearly three times as often. Then it found that changing the send interval moved inbox placement more than any copy change did. That combination is the actual story.

STRATEGYCOLD EMAIL

Everybody I talk to about AI SDR performance wants to talk about the copy. Is the personalization good enough. Does it sound human. Should we run it through another model to take the edges off. Fair questions. Wrong questions, mostly.

A comparison published this year put 100,000 cold emails side by side, 50,000 written by AI and 50,000 written by people, across a window running from October 2025 into April 2026. The analysis of AI SDR versus human-written performance reports the AI side replying at 4.1 percent against 5.2 percent for humans. That is a gap. It is not a catastrophe, and if you stopped reading there you would conclude the machines are close and getting closer.

Keep reading and it stops being a copy story.

4.1% vs 5.2%
reply rate, AI-written against human-written, across 100,000 paired cold emails
0.7% vs 1.1%
meetings booked, AI against human, in the same comparison
8% vs 3%
spam flag rate, AI against human, nearly three times higher for the AI side
6%
bounce rate, identical on both sides, which is what makes the rest comparable
THE SHORT VERSIONThe reply gap is small. The meeting gap is large. The spam gap is enormous. Those three facts in that order tell you the AI side is not losing on sentences, it is losing on trust, and trust damage compounds the further down the funnel you look. Meanwhile the single biggest lever measured in the whole dataset was the send interval, which no copy tool touches.

What the 100,000 email comparison actually found

Start with the bounce rate, because it is the least interesting number and the most important one. Six percent on both sides. Identical.

That means the two groups were working from list data of the same quality. Nobody handed the humans a clean list and the machines a scraped one. Whatever separates the two groups downstream, it is not data hygiene, and that single control is what makes the rest of the numbers worth arguing about.

METRICAI-WRITTENHUMAN-WRITTENRELATIVE GAP
Bounce rate6%6%None. List quality matched
Reply rate4.1%5.2%Humans 27% higher
Positive replies1.4%2.1%Humans 50% higher
Meetings booked0.7%1.1%Humans 57% higher
Spam flag rate8%3%AI 2.7x higher

Look at the third column going down. Twenty seven percent. Fifty percent. Fifty seven percent. The gap widens at every stage of the funnel, and the relative-gap column is our arithmetic on the reported figures, not a separate claim from the study.

That shape matters. If AI copy were simply a bit blander than human copy, you would expect a roughly constant penalty all the way down. A slightly worse email gets slightly fewer replies, and those replies convert at about the same rate. That is not what happened. The AI-generated emails did not just get fewer replies, they got worse replies, and the replies they did get turned into meetings less often.

Something is filtering for quality of intent, not quality of prose. And the spam number tells you what.

AI SDR performance is losing on trust signals, not sentences

Eight percent of the AI-written emails got flagged as spam. Three percent of the human ones did. Same lists, same bounce rate, nearly three times the flag rate.

Spam filters are not literary critics. They are not sitting there deciding your second paragraph reads like it was generated. They are looking at sending patterns, domain reputation, volume ramps, engagement history, and how closely a message resembles other messages hitting the same infrastructure. That is a behavioural fingerprint, and AI-assisted outbound has a very distinctive one, because the whole point of automating it is to do more of it, faster, from more addresses.

So the copy is a red herring. The machine is not being punished for how it writes. It is being punished for how it operates: higher volume, tighter intervals, more templated structure across a shared sending footprint, and a reply signal too thin to build a positive reputation on. Once you land in a filter, the recipients you needed most never see the message, and those are disproportionately the ones at larger companies with the strictest mail policies. Which is exactly where your meetings were going to come from.

01The penalty is infrastructural, so the fix is infrastructuralYou cannot prompt your way out of a domain reputation problem. Separate sending domains by campaign type, keep volumes per mailbox low enough to look like a person, and warm properly before scale. This is the same argument as the one about how many sending domains you actually need, which almost everyone gets backwards.
02Thin reply signal is self-reinforcingFilters use engagement to decide who to trust. Low reply rates lower your reputation, which lowers deliverability, which lowers reply rates. AI outbound tends to enter that loop faster because it scales volume before it has earned any positive signal to scale against.
03Template similarity is measurable across sendersIt is not just your own repetition that matters. If a widely-used tool generates structurally similar messages for thousands of customers, the pattern is visible at the filter's scale even when it is invisible at yours. Being one of ten thousand accounts producing the same shape of message is a reputation input you do not control and should assume is working against you.
04The meeting gap is the number to reportReply rate is the vanity metric of outbound. A 27 percent reply gap sounds survivable and a 57 percent meeting gap does not, and they came from the same test. Report the one that maps to pipeline, or you will keep approving programs on the flattering half of the data.

None of this says do not use AI in outbound. It says the thing AI is good at, which is producing volume, is the thing that is most expensive in this channel right now. That is a genuinely awkward finding for the category and it deserves to be said plainly rather than buried. We made a version of this argument about what sender credibility now costs you, and this dataset is the operational evidence for it.

Cadence beat every copy variable in the test

Here is the part that should have been the headline.

Moving the send interval from one day to three days took inbox placement from 71 percent to 93 percent. Twenty two percentage points. No new copy, no new personalization layer, no model upgrade. Just waiting longer between touches.

One day between sends71%
Three days between sends93%

Inbox placement by send interval, as reported in the 100,000 email analysis

Sit with the scale of that for a second. Twenty two points of inbox placement is the difference between a sequence that works and a sequence that quietly does not, and it dwarfs any plausible lift from rewriting a subject line. The study's own framing is that cadence beats content, and on these numbers that is hard to argue with.

It is also the least glamorous finding imaginable, which is probably why nobody is selling it. There is no product in wait longer. There is no demo. You cannot put a slider on it. So the entire conversation drifts back to copy, where the tooling is, and away from the setting that actually moved the number.

The lever that moved most was the one nobody can sell you. That is usually how it goes.

The mechanism is not mysterious. Tighter intervals concentrate volume, and concentrated volume from a young or lightly-warmed domain is one of the loudest signals a filter has. Spreading the same sequence over more days lowers the peak rate per mailbox per hour, gives positive engagement time to register between touches, and stops the pattern from looking like a burst. It is the same physics behind monitoring your sending reputation directly rather than inferring it instead of guessing from open rates that were never reliable anyway.

How to read a vendor-adjacent benchmark without getting played

Now the caveats, because I would rather you take this seriously than take it literally.

This is not a peer-reviewed study. It is an analysis compiled by an agency from aggregated data across several outbound platforms plus their own databases. That is a legitimate way to build a dataset and it is also a way to build one with its thumb on the scale, because the people who assemble these numbers usually have a service to sell on the back of them. Take the directions seriously and hold the decimal places loosely.

Three specific things I would want before treating any of these figures as settled. First, how were the AI and human groups assigned, because if the human-written emails went to warmer or better-researched accounts then the whole comparison collapses. Second, what counts as a positive reply, since that definition swings by a factor of two between vendors. Third, whether the cadence finding was tested inside both groups or only observed across them, because a correlation between slower sending and better placement could partly be a proxy for more careful operators.

Direction over magnitudeThat AI-written outbound gets flagged more and books fewer meetings is consistent with what the deliverability community has been reporting for two years. Believe the direction. Do not quote 2.7x in a board deck as though it were measured in your account.
Replicate the cadence test yourself, cheaplyThis is the rare finding you can verify in three weeks. Same list segment, same copy, two send intervals, measure placement with a seed set rather than opens. If you get anything close to twenty points, you have found more value than a quarter of copy testing would have produced.
Ask who is not in the sampleAggregated platform data is data about people who use those platforms. The best outbound teams I know send less, from fewer mailboxes, with more research per account, and they are systematically underrepresented in any benchmark built this way. The ceiling in these datasets is not the real ceiling.

That is not scepticism for its own sake. It is that outbound benchmarks get quoted for years after the conditions that produced them have changed, and the ones that spread fastest are the ones with the most quotable single number. The 4.1 against 5.2 will travel. The 71 to 93 will not, and it is the one worth having.

The AI SDR performance fixes worth doing this month

Four things, roughly in order of how much they will move.

Change your send interval first. Go from one day to three between touches on one live sequence, hold everything else constant, and measure inbox placement with seed accounts rather than open tracking. This is a settings change. It costs you nothing but calendar time, and on the evidence above it is the highest-leverage thing available to you this month.

Then split your reporting so the meeting rate leads and the reply rate follows. If your outbound dashboard opens on reply rate, the widening funnel gap this data describes is invisible to you, and you will keep making decisions on the number that flatters the program. The gap between a 27 percent difference and a 57 percent difference is exactly the gap between renew this and kill this.

Third, audit where AI is actually sitting in your process. Using it to research an account, summarize a trigger, or draft a first pass a human then rewrites is a different thing from using it to generate and send at volume. The first shows up as leverage. The second shows up as an eight percent spam flag rate. Most teams describe themselves as doing the first and, when you look at the sending logs, are doing the second.

Fourth, put your domain and mailbox architecture on the same review cadence as your copy. Sending domains, per-mailbox volume, warmup state, and reputation monitoring get reviewed monthly in the cold email programs we run, because these are the variables that decide whether any of the copy work gets seen at all. For tech and SaaS teams selling into enterprises with strict mail policy, this is not a marginal concern, it is most of the game.

You are not behind because your AI copy is not good enough yet. That framing has cost a lot of teams a lot of quarters. The copy is fine. The volume it enables, sent on the infrastructure most people are sending it on, at the intervals most people are using, is what is costing you meetings. Fix the operating envelope first. Then, if you still want to argue about the second paragraph, at least somebody will be reading it.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

TT
Tyler TruffiMANAGING PARTNER, SOMETHING INC.

Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.