Every debate in cold email eventually reduces to the same underlying question: is the reply rate gap between average and elite senders a writing problem or an infrastructure problem. Two of the loudest voices in the space right now are, in effect, arguing opposite answers, and both are right about the specific piece of the problem they're each actually looking at.
The craft case: poke the bear
Josh Braun's approach, taught through his newsletter, cold call scripts, and a large body of shared LinkedIn breakdowns, starts from a specific diagnosis: most cold outreach fails because it pitches before it's earned the right to. "Poke the bear" is his term for a different opening move, a single, low-pressure question that surfaces a problem the prospect hasn't consciously named yet, rather than a claim about your product's features. A representative example from his own material: "I often hear that for orgs your size, it takes a few days to calculate and send statements to reps. How does that compare with your experience?" That's not a pitch. It's an invitation to reflect on a specific, plausible pain point, phrased as a question rather than an assertion.
The mechanism behind why this works, in Braun's framing, is that brains are wired to pay attention to problems, not solutions. Leading with a claim about your product asks a stranger to evaluate a sales pitch. Leading with a sharp, specific question about a problem asks them to evaluate their own situation, which is a much lower-friction ask and doesn't trigger the reflexive skepticism a pitch does. "Ditch the pitch" is the shorthand version of the whole philosophy: get to the problem fast, because problems get attention, and let the prospect draw their own conclusion about whether they need to talk to you next.
The craft case is genuinely hard to scale, and that's not a knock on it, it's the honest tradeoff. Writing a poke-the-bear question that's specific enough to land and general enough to apply to a real segment of your list requires real understanding of that segment's actual operational pain, not a mail-merge field. Braun's own material is built almost entirely around teaching that skill directly to individual reps, one script, one call breakdown, one LinkedIn post at a time, which is a craft-transmission model, not a systems one.
There's a reason this style of teaching resonates as widely as it does: it's legible in a way infrastructure advice usually isn't. A rep can read one script breakdown, understand exactly why the wording works, and apply the underlying pattern to their own accounts the same afternoon. That immediacy is a real advantage over advice that requires a tooling build-out before it produces a single sent email. It's also, on its own, a ceiling. One rep, however skilled, can only write and personalize so many genuinely sharp questions in a day, and that ceiling doesn't move no matter how good the underlying skill gets.
The system case: let AI handle the research
Instantly's 2026 cold email benchmark report, drawn from billions of tracked emails, makes a different argument about where the gap between average and elite senders actually comes from. Its framing: elite cold email teams run intelligence-led outbound, and increasingly use AI agents that handle roughly 80% of the research and sequencing work that used to consume a rep's entire day. The lever, in this telling, isn't a sharper individual message, it's the ability to do genuine, specific research on every single prospect at a volume no human researcher could sustain manually, then sequence outreach around whatever that research actually surfaces.
This is the infrastructure answer to the same reply-rate gap, and it has its own real data behind it. Signal-anchored campaigns, ones built around a specific, timely trigger rather than a generic list, post reply rates in the 15 to 25% range, roughly five times the blended average. Getting there at any meaningful volume requires research depth per prospect that simply doesn't scale through a human doing it manually one account at a time, which is exactly the gap AI-agent-assisted research is built to close: not writing a smarter message, but surfacing a real, specific, current signal about a specific account fast enough to act on it before it goes stale.
The system case has its own honest tradeoff, mirroring the craft case's ceiling in the opposite direction. Research depth at scale is only valuable if something is actually done with it, and an AI agent that surfaces a genuinely sharp, specific signal about a prospect still needs that signal translated into an opening line a human would actually want to read. Handing 80% of research and sequencing to automation doesn't automatically hand over the remaining 20% well; it just changes what the remaining human effort needs to be spent on. A team that treats the automated research output as a finished message rather than raw material for one is trading a manual bottleneck for an automated version of the same generic-outreach problem, just at higher volume.
| DIMENSION | POKE THE BEAR (CRAFT) | AI-AGENT RESEARCH (SYSTEM) |
|---|---|---|
| What it optimizes | The specific wording and framing of the opening message | The volume and freshness of the signal the message is built around |
| Scales via | Teaching the skill to more humans, one at a time | Automating research and sequencing across an entire list |
| Fails when | A sharp question gets sent against a stale or irrelevant signal | A perfectly-researched signal gets wrapped in a generic, template-sounding message |
| Primary data point | Qualitative: reply quality and conversation depth per send | Quantitative: reply rate lift across large, tracked send volumes |
Where they actually agree
Read past the framing and both camps are making the identical underlying argument: generic outreach is dead, and specificity is the entire game. Braun's version of specificity lives in the question's wording. Instantly's version lives in the research feeding that question. Neither camp is actually arguing for spray-and-pray at scale, and neither is arguing that craft or research alone, without the other, is sufficient. Jordan Crawford's pain-qualified segments framework, which Something Inc. has already covered in depth, sits almost exactly at the intersection of these two positions: use data to find real, current pain at scale, the systems side of the argument, then let that pain define the message itself rather than generic personalization, which is functionally the same instruction as "poke the bear," just arrived at from the infrastructure side instead of the messaging side.
The reply-rate numbers make the same point from a different angle. A 3.43% blended average and a 15-25% signal-anchored range aren't two different tactics competing for credit, they're two ends of the same spectrum, and both Braun's and Instantly's prescriptions push in the identical direction: toward more relevance per send, whether that relevance gets built through a sharper question or a deeper research pass. The disagreement, to the extent there is one, is about which half of the relevance problem is currently the more under-invested lever for a given team, not about whether relevance itself is the goal.
The verdict: sequence, don't choose
For a team actually building an outbound motion right now, the actually honest answer here is that these aren't really competing philosophies to pick between at all, they're sequential layers, and most genuinely underperforming outbound programs are missing one of the two entirely, not both at once. A team with strong AI-assisted research and a generic template wrapped around it is leaving the Instantly half of this argument on the table: it has the signal but not the craft to use it well. A team with genuinely sharp, well-crafted messaging sent against a stale or generic list is leaving the Braun half on the table: it has the craft but not the research depth to aim it correctly.
It's worth resisting the temptation to treat this as a 50/50 split by default, too. The right ratio between craft investment and research-system investment depends on where your program actually sits today, and that starting point varies enormously between a two-person founder-led motion and a twenty-rep outbound team with an established tech stack already in place. Audit which half your program is actually missing before investing further in either. If your reply rates are flat despite genuinely researched, current signals feeding your sends, the problem is probably in how those signals get translated into an opening line, and Braun's material is the more useful next investment. If your messaging already reads sharp and specific but your signal-to-noise ratio on targeting is weak, more AI-driven research depth per account, the direction Instantly's data points toward, is the higher-leverage fix. Most cold email programs Something Inc. has audited are missing the research-depth half more often than the craft half, which tracks with why our own reply-rate benchmark data still shows such a wide spread between vendors measuring functionally the same channel: the underlying list and signal quality varies far more than most teams' messaging polish does.
Build the measurement to actually tell the two apart, since reply rate alone doesn't isolate which half is broken. A B2B team running this audit honestly, segmenting reply data by signal freshness and by message specificity separately rather than blending both into one number, will generally find one lever is doing far more work than the other, and that's the lever worth doubling down on first, regardless of which camp's framing got you to look.
The teams that actually post the double-digit numbers both camps are implicitly describing aren't the ones that picked a side in this debate. They're the ones that treated craft and system as two separate line items on the same roadmap, funded both, and kept measuring each one honestly enough to know which one needed the next dollar. That's a less satisfying answer than "the real secret is X," but it's the one the data from both camps actually supports once you stop reading them as competing claims and start reading them as two halves of the same argument.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.