Somebody on your team is reporting a 62% open rate this week and drawing conclusions from it. Subject line B beat subject line A. Tuesday beats Thursday. The 9am send is working. None of that is safe to believe, and the reason has nothing to do with your copy.
Apple shipped Mail Privacy Protection in September 2021. When a device is on Wi-Fi and charging, Apple routes the message through its own proxy servers and pre-loads the full contents — every image, including the invisible one-pixel image your sending tool uses to register an open. The pixel fires. Nobody read anything.
Who is actually opening your email
Scale decides whether this is a rounding error or a broken metric, and the scale is not subtle. Litmus Email Analytics put Apple Mail at 51.52% of the global email client market in January 2026. Bird's analysis found more than 55% of all global opens coming from Apple devices running MPP.
Share of the global email client market (Litmus Email Analytics, January 2026)
Half your list is behind a system that fires the pixel on delivery. That is not a measurement error you can correct with a coefficient, because you cannot tell which half of any given campaign was proxied. The metric is not noisy. It is structurally uninterpretable.
It gets worse for the tests people actually run on it. A subject line A/B test assumes the only thing differing between arms is the subject line, and that opens respond to it. But MPP fires regardless of what the subject says — the proxy does not read. So roughly half of every arm is a constant, diluting whatever real effect exists toward zero. You can run that test for months, watch it produce a 2% difference, and ship the losing variant with full confidence.
The same contamination sits under send-time optimization. If a meaningful share of opens register at delivery rather than at reading, your open-by-hour curve is partly a picture of when your own tool sent the mail. Teams have rebuilt entire sending schedules around that artifact.
The three false positives inside every open rate
That third one is what kills the salvage attempts. If the bias were purely inflationary you could at least track the trend and ignore the level. But with proxies inflating and image blocking deflating, on populations that differ by segment and by campaign, even the direction of a week-over-week change is unreliable.
Open tracking pixels vs reply tracking: what each one proves
| SIGNAL | WHAT HAS TO HAPPEN | WHO CAN FAKE IT |
|---|---|---|
| Open | An image loads | Apple's proxy, a security scanner, a preview pane |
| Click | A URL is requested | A link-scanning gateway, a spam filter, a preview crawler |
| Reply | A human writes and sends words | Out-of-office autoresponders, and that is the whole list |
Replies are the only row where a person had to decide something. Filter out-of-office responses — every decent tool does this now — and what remains is a human who read your message and chose to spend thirty seconds on you. It is a smaller number than your open rate and it is the only one that has ever predicted pipeline.
Clicks sit in an awkward middle and are worth a specific warning. They feel more solid than opens because a URL request seems more deliberate than an image load, but link-scanning gateways follow every link in an inbound message as a security measure, often within seconds of delivery. In enterprise segments that produces click events on messages no human has opened, and because the scan happens once per message you can get a click rate that looks like genuine engagement from an account that never saw the email. If you must keep one of the two machine signals, keep neither.
“An open is evidence that a machine touched your email. A reply is evidence that a person considered it. Only one of those is a marketing metric.”
Instantly's 2026 benchmark puts the average cold email reply rate at 3.43%, with top campaigns clearing 10%. Those are small numbers next to a 62% open rate, and that is precisely why teams cling to opens: the vanity metric is the one that makes the weekly report look survivable. It is also why cold email reply rate benchmarks are the only benchmarks worth holding a campaign against.
The cost nobody prices in
Here is the part that turns this from a measurement argument into an operational one. The tracking pixel is not free to carry. It is a remote image load from a third-party domain embedded in a one-to-one business email — a pattern that looks far more like bulk marketing than like a note from a person, and filters weight it accordingly. GlockApps' deliverability research flags tracking pixels as a placement risk for exactly this reason.
So the trade is worse than it looks. You are accepting a measurable deliverability penalty in order to collect a number that more than half the market has already invalidated. If you are running warmed domains and a carefully staged ramp, the pixel is quietly working against all of it — the same way an unmanaged sending pattern undoes a warmup pool you spent six weeks building.
What to instrument instead
Dropping opens does not mean flying blind. It means moving the diagnostic layer to signals that survive contact with modern mail infrastructure.
That third play is the one people miss. Open rate was doing double duty — it was a message-quality signal and an early warning that mail had stopped landing. Removing it without replacing the second job is how teams end up discovering a placement collapse three weeks late, which is the failure mode behind most of the Gmail spam threshold incidents we get called into.
Turning it off without going blind
Expect the first two weeks to feel worse. Your dashboard will lose its biggest number and the remaining ones are small. Somebody will ask whether performance dropped. It did not — you stopped counting proxy servers as prospects, and the reply rate you are now looking at was the real performance the whole time.
The teams that make this switch cleanly do one thing first: they run a fixed seed-list placement test before turning the pixel off, then again two weeks after, so they can show the deliverability gain in the same meeting where they explain the missing metric. That reframes the change from losing data to trading a fake number for inbox placement, which is an argument that survives contact with a sceptical stakeholder. It is how we sequence it on cold email engagements.
See where you are cited today
A free snapshot audit of your rankings and AI citations before we ever talk.
Josh leads work at the intersection of SEO and generative engines at Something Inc., helping B2B brands get ranked and cited across every major AI engine.