Something Inc.Schedule a free consultation
STRATEGY

The AI wrote your cold emails. Your domain is paying for it.

In a 100,000-email paired analysis, AI-written and human-written cold emails bounced at exactly the same rate — and AI copy got flagged as spam nearly three times as often. The gap is not deliverability. It is what the copy sounds like.

TTTyler TruffiManaging Partner · AUG 13, 2026 · 11 MIN READ

Every outbound team is asking the same question this year, usually in a Slack thread that goes quiet without resolving: do we let the AI write the emails? The arguments run on vibes in both directions. There is now a paired dataset that answers it with numbers, and the answer is more specific and more useful than either camp expected.

The Digital Applied team published an analysis on April 26, 2026 covering 100,000 paired cold emails — 50,000 AI-generated, 50,000 human-written — matched on persona, ICP firmographics, sequence stage, sender domain age and sender domain authority, sent between October 2025 and April 2026. One caveat up front, because it matters: the methodology describes the data as sourced from Smartlead, Instantly, Apollo and proprietary aggregated sources. That is careful aggregation rather than a controlled first-party experiment, and the matching does a lot of work. Read the deltas, not the absolute levels.

6% / 6%
bounce rate, AI versus human — identical
8% vs 3%
spam-flag rate, AI versus human
4.1% vs 5.2%
reply rate, AI versus human

The question every outbound team is asking

The framing in most of these debates is wrong from the start. Teams argue about whether AI copy is good, as if quality were one dial. The data says AI-written outbound fails in a specific place, succeeds in a specific place, and is close to neutral everywhere else. Once you know which is which, the decision stops being ideological and becomes a routing problem: which segments get machine-written first drafts, which get a human, and what you do about the sending infrastructure either way.

METRICAI-GENERATEDHUMAN-WRITTENWHAT THE GAP TELLS YOU
Reply rate4.1%5.2%Roughly a point of difference — real, but smaller than either camp claims
Positive reply rate1.4%2.1%The gap widens on quality of reply, not just volume
Meetings booked0.7%1.1%Compounds down the funnel; a 1.6x difference at the outcome that pays
Bounce rate6%6%Identical. Copy has nothing to do with whether an address exists
Spam-flag rate8%3%The headline finding. 2.7x, and it lands on your sending domain

Two of those rows do the analytical work. Bounce is identical, and spam-flag is nearly triple. Digital Applied's own line on this is the cleanest summary anyone has written of a distinction most teams never make: bounce rate is a list problem, spam-flag is a content problem. Those two failure modes get lumped together as deliverability, they have completely different causes, and only one of them is affected by who or what wrote the email.

AI cold email vs human: what the paired numbers show

Start with the reply gap, because it is smaller than the discourse suggests. A point and a bit of raw reply rate separates the two. If you are running high volume into a large addressable market and you have the deliverability headroom, that gap is survivable — arguably it is the price of covering ten times the accounts.

But follow it down the funnel. Positive replies run 1.4% against 2.1%, and meetings booked run 0.7% against 1.1%. The gap widens at every stage, which is what you would expect if AI copy is slightly more likely to produce a polite non-answer and slightly less likely to produce the reply that starts a conversation. By the time it reaches the metric your revenue team cares about, a modest copy difference has become a 1.6x difference in booked meetings. Volume can close that. Volume also raises the number that matters most in this dataset.

One more piece of trend context. The same analysis puts human performance flat year over year at 5.2%, while AI-written performance improved by 1.3 percentage points from 2.8% in 2024. The machine side is closing, slowly, from below. Anyone extrapolating that line to a crossover date is guessing; anyone assuming the current gap is permanent is guessing in the other direction.

Identical bounce, triple the spam flags

The 8% versus 3% spam-flag split is the finding to act on, because a spam complaint is not a lost email. It is damage to the asset every future campaign depends on. Gmail's own guidance treats complaint rate as a hard operating constraint, and Google's Postmaster tooling now surfaces a plain-language verdict when a sender's spam rate goes above 0.1 percent. On the Microsoft side, domains sending 5,000 or more messages a day to Outlook, Hotmail, Live or MSN addresses must pass SPF and DKIM with DMARC alignment or get rejected outright with a 550 5.7.515 or 550 5.7.509 response — not junked, rejected.

Those thresholds are the reason the spam-flag number is worse than it looks. Reply rate is a per-campaign outcome you can recover from next week. Complaint rate is a rolling reputational signal attached to a domain and an IP, and once it degrades, every campaign that follows starts from a worse position — including the ones a human wrote.

THE ASYMMETRY TO HOLD ON TOA reply-rate loss costs you this quarter's meetings. A spam-flag increase costs you the channel. Treat them as different classes of risk, and never trade the second for the first.
Bounce rate is a list problem. Spam-flag is a content problem. Most teams run one dashboard for both and wonder why the fixes never work.

Why does machine-written copy draw more complaints? The dataset does not say, and neither will we — that would be inventing a mechanism to fit a number. What we can say is that the fix follows the diagnosis regardless of cause: if spam-flag is a content problem, the intervention belongs in review of the copy, not in more warm-up or another sending domain. Buying more infrastructure to absorb a content problem is the most expensive mistake in outbound, and it is the default response.

The variable that beat copy entirely

Here is the part that should reorganize your priorities. Digital Applied's own conclusion, stated plainly, is that the single most impactful variable in the dataset is not the subject line, body length, personalization token or whether a human or a machine wrote it. It is the interval between sends.

Their supporting figure: on a one-day cadence, inbox placement ran 71% for AI-written and 86% for human-written sends. Cadence is a setting. It takes one person one minute to change, it costs nothing, and in this dataset it moved placement further than the entire AI-versus-human argument that has consumed the last eighteen months of outbound discourse.

Human-written, one-day cadence86%
AI-written, one-day cadence71%

Inbox placement on a one-day send cadence, AI-written versus human-written (Digital Applied, 100K paired emails, Oct 2025-Apr 2026)

Read that chart with the vendor guidance next to it. Austin Hughes, co-founder and CEO of Unify, published operating numbers in July recommending roughly three weeks of mailbox warm-up, a default of 25 emails per day per mailbox configurable up to 65, and sequences of four to seven emails. Twenty-five a day is materially more conservative than the fifty most agencies run, and it comes from someone whose business grows when you send more.

Where AI cold email vs human flips by vertical

The aggregate hides a reversal. In the same dataset, AI-written email replied at 6.1% in SaaS — above the human average — and at 1.9% in financial services. Digital Applied's summary of it is blunt: in SaaS, AI cold email beats human; in financial services, it gets you blocked.

That spread is a routing rule waiting to be written. Technical and software buyers appear to tolerate, or possibly prefer, the register machine-written outbound produces. Regulated, relationship-led buying committees do not, and the cost of finding out shows up in your complaint rate rather than in a polite decline. If your ICP spans both, a single blanket policy on AI-written copy is guaranteed to be wrong for half your pipeline.

1Segment the policy by vertical, not by team preferenceSoftware and technical buyers are the defensible place to run machine-written first drafts at volume. Regulated and relationship-led segments get human copy, or at minimum human review before send. Write this down as a rule with named owners rather than leaving it to whoever configured the sequence.
2Isolate sending infrastructure by risk tierIf you are going to run AI-written volume, do not run it from the same domains carrying your highest-value human sequences. Separate domains mean a complaint-rate problem in one program cannot take the other down with it.
3Fix cadence before you touch copyThe interval between sends outperformed every copy variable in this dataset. Slow the cadence, respect a per-mailbox daily ceiling in the twenties rather than the fifties, and re-measure before spending a sprint on subject lines.
4Alert on complaint rate, not just on repliesGoogle's Postmaster data now returns an explicit verdict when spam rate goes above 0.1 percent, and it is machine readable. Wire that into the same alerting that watches your bounce rate, because by the time reply rate reflects a reputation problem you have been burning the domain for weeks.
5Review AI copy for the complaint driver, not for qualityThe reply gap is a point. The complaint gap is nearly triple. Human review of machine drafts should be looking specifically at what makes a recipient press the button, which is a narrower and more tractable editing job than making the email good.

What the vendor numbers claim, and how to read them

For contrast, the vendor side of this market publishes a very different picture. AiSDR, in a study covering 75 companies across SaaS, fintech and IT services, reports reply rates climbing from a 2.4% baseline to 6.8% by month three and 8.2% by month six, meetings per month rising from 12 to 38, output per SDR growing 368%, and 78% of teams cutting SDR headcount by around 30%.

Those numbers are not necessarily false, and they are not evidence either. They come from a vendor measuring its own product, with no sampling methodology beyond a company count, describing what the write-up itself calls real team experiences. They also contradict the paired analysis directly on the one question both purport to answer. When two datasets disagree that sharply, the tiebreaker is method: matched pairs across a hundred thousand sends beats 75 self-selected customers, even after you discount the aggregation caveat.

This is the same discipline we applied to the 18x spread across published reply-rate benchmarks earlier today, and it is the whole job when a category is this new: ask what was counted, who counted it, and what happens to the vendor if the number comes out low.

Do this next

Pull your complaint rate by sending domain for the last ninety days and put it next to your reply rate. If complaints are climbing while replies hold steady, you are trading the channel for this quarter's number and the trade will come due. Then slow your cadence and cap daily volume per mailbox before you change a single subject line, because that is where the evidence says the leverage is.

After that, split your policy by vertical and separate the infrastructure. Our cold email engagements run machine-assisted drafting inside exactly these guardrails, our tech and SaaS work is where the machine-written approach earns its place, and we have written on auditing an AI agent in your outbound stack and what skipping warm-up actually costs. Digital Applied's full paired analysis has the per-vertical tables if you want to check the segments against your own ICP.

See where you are cited today

A free snapshot audit of your rankings and AI citations before we ever talk.

TT
Tyler TruffiMANAGING PARTNER, SOMETHING INC.

Tyler leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.

Free consultation

Let us be the last SEO agency you ever work with

A 30 minute call and a free audit of your SEO and GEO position. You keep the findings either way.