Questions about this study
State of Email Outreach: which wording correlates with opportunities
12 questions answeredMethod challenges included
What did this study actually find?
It looked at 287,790 cold emails sent between 2026-01-26 and 2026-08-14 and checked whether the email got a reply. Emails were also tagged by content segment, such as whether they quoted a percentage statistic, a named client, a case study number, and so on. The email segment that quoted a percentage statistic replied at 0.8%, the lowest reply rate of the six content segments measured. Every other segment, including emails with no proof element at all at 1.3%, replied more often. The overall reply rate across all emails was 1.2%.
Is a difference that small even meaningful?
It's fair to be skeptical here. The six segments range from 0.8% to 1.8%, which is a spread measured in fractions of a percentage point, not a doubling or halving of reply rates. The percentage-stat segment is consistently at the bottom, which is the finding, but nobody should read this as 'quoting a stat kills your reply rate.' It's a small, consistent disadvantage, not a dramatic one.
Does this mean including a percentage in a cold email causes fewer replies?
Not on its own. This is an observational comparison, not a controlled experiment: emails were not randomly assigned to include or exclude a percentage, so the people or teams who chose to quote a stat may differ in other ways from those who didn't, for example in industry, offer type, or list quality. To call it causal you'd want an A/B test where the only thing that changes between two otherwise identical emails is the presence of a percentage. This study can tell you the association is there and roughly how consistent it is, not why.
Why would quoting a percentage hurt reply rates instead of helping?
The study doesn't measure the mechanism directly, but a plausible explanation is that a percentage statistic reads as a sales claim, which triggers the same pattern recognition a reader uses to skip marketing copy. Recipients see a lot of cold email, and a specific-sounding number without context can signal 'template' rather than 'someone read my situation.' An alternative explanation is that percentage stats get used disproportionately in higher-volume, more generic campaigns, which would produce the same pattern without the number itself being the cause. This study can't distinguish those two on its own.
How rare is a reply, overall?
Rare. Across all 287,790 emails, only 3,368 got a reply, which is 1.2%. The remaining 284,422 (98.8%) got no reply. Because replies are such a small minority of outcomes, every segment comparison in this study is really a comparison between small minorities, for example 0.8% versus 1.3%, and small absolute differences can look more dramatic when both numbers are already close to zero.
Could this just be one bad sending list dragging down the percentage-stat segment?
That's a real concern, and the data lets us check it partially. Across the 45 sending groups measured, the replied class ranged from 0% to 4.5%, which is a much wider spread than the difference between content segments. That means which list or campaign an email came from moved reply rates more than which content segment it fell into. The largest single sending group only accounted for 14.8% of all emails, so no one list dominates the total, but it does mean list-level variation is a bigger factor than the content tag, and some of that variation could be correlated with which segment a list tends to use.
What counts as a 'percentage statistic' in this study versus, say, a case study number?
A percentage statistic is a claim phrased as a percentage or rate, such as an efficiency or lift figure written with a percent sign. A case study number is a different kind of quantified claim, typically a count or specific figure tied to a named example rather than expressed as a rate. The study segmented 79,293 emails as percentage-stat and 17,738 as case-study-number, and each replied at a different rate (0.8% and 1.5% respectively), so the specific form the number takes appears to matter, not just the presence of any number.
What's the 'no proof' baseline, and why does it matter?
No proof means the email contained none of the proof elements tracked in this study, no named client, case study number, timeframe stat, social proof line, or percentage statistic. It's the closest thing to a control group here. It's also the largest segment by far, 171,160 of the 287,790 total emails. It replied at 1.3%, close to the overall average of 1.2%, which is exactly the baseline the percentage-stat segment fell below. That comparison is what makes the percentage-stat result notable rather than just noise.
Should I stop putting percentages in cold emails based on this?
The direction of the finding says a percentage statistic is the weakest-performing proof element measured here, not that it actively repels replies in a large way. Given the gap is small and the study is observational, a reasonable action is to test it directly on your own list, sending otherwise identical emails with and without a percentage claim, rather than removing every stat from your templates on the strength of this alone. If you do keep percentage claims, this data would suggest pairing them with something else, since no-proof and other proof types all slightly outperformed percentage-only claims here.
Did the researchers exclude any emails or clean the data before running this?
No. The fact table for this study reports 0 excluded rows and 0 mitigations applied, meaning all 287,790 rows collected were used as sent, with no filtering, deduplication, or adjustment. That's worth knowing both ways: it means the result isn't shaped by a judgment call about what to throw out, but it also means any duplicate sends, bounces misclassified as no-reply, or other data quality issues in the raw logs would be sitting in these numbers unaddressed.
Is one segment ever counted in more than one category, like an email with both a percentage and a named client?
The study doesn't report how segment assignment handled emails with multiple proof elements, so that isn't something this data can answer. What we do know is each segment's own count: 79,293 emails in the percentage-stat segment, 10,958 in named-client, and so on, summing close to but not necessarily equal to the full 287,790 depending on overlap rules. If you need to know whether segments are mutually exclusive for your own analysis, that would require going back to the original tagging methodology, which isn't included here.
How consistent is the 'no reply' outcome across different sending groups?
Very consistent at the high level, less so in the details. Across the 33 groups measured, the no-reply share ranged from 95.5% to 100%, with a median of 99.2%. So while no-reply is always the overwhelming majority outcome, individual groups still vary by several points, which lines up with the earlier point that which group an email belongs to explains more variation than which content segment it falls into.
Free strategy session
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
Omar JenblatFounder & CEO, BusySeed
- Your category measured the same way
- Your own numbers, not a sample deck
- Fifteen minutes, no obligation
