Methodology
State of Email Outreach: which wording correlates with opportunities
Engines and exact versions
A study that does not name the model version it ran against is not reproducible, because the answer changes when the model does.
| Engine | Model version |
|---|
Run window, UTC: [object Object].
Method
Unit of analysis. One unit is one sent cold email, matched to whether it received a reply within the measurement window. There were 287,790 such units (287,790 rows, with 0 excluded, so every sent email that met the study's inclusion criteria is represented). An email is the right unit here because reply behavior is a property of the individual message a recipient actually reads, not of the campaign or the sender; aggregating at the campaign level would hide variation in how a given template was worded from send to send.
Outcome classes. Every unit falls into exactly one of 2 classes: replied or no reply. This is a binary, unambiguous outcome, read directly from whether a reply was recorded, not inferred from opens or clicks, which removes a layer of interpretation that plagues weaker email studies (an open does not mean a read, and a click does not mean interest).
Content segments. Each email was also coded into one of 6 content segments describing the proof element in its body: percentage_stat, named_client, case_study_number, timeframe_stat, social_proof, or no_proof. These are not mutually reinforcing categories layered on top of each other; each email was placed in the segment that best described its dominant proof element, which is what allows a clean comparison of reply rate by segment.
Groups. Beyond content segment, the data also carries 45 distinct sending groups, likely representing different campaigns, senders, or send batches. The largest single group accounts for 14.8% of all units, meaning no one group dominates the dataset. This matters for the next section: a finding that holds only because one huge, unusual group skews the average is a much weaker finding than one that holds across many groups of comparable size.
What was not done. No mitigations were applied to the raw counts (0), and no rows were excluded (0). The numbers reported are the numbers observed, not adjusted, reweighted, or filtered for outliers. That is a deliberate choice for a first-pass descriptive study: adjustments can correct for known biases, but they can also quietly launder a result. Reporting the raw observed rates keeps the finding auditable.
Limitations we volunteer
Written by us, before anyone else found them.
- Single pass. Run-to-run variance is not characterised.
- Gemini's cited sources are largely unavailable through Google's API, so source analysis rests on the other engines.
Terms used in this study
- Segment
- A grouping of emails by a content feature they share, such as containing a percentage statistic, a named client, or no proof element at all. This study measured six segments.
- Class
- One of the two possible outcomes recorded for each email: replied or no reply. Every email in the dataset falls into exactly one class.
- Sending group
- A distinct batch or campaign of emails, likely differing by list, sender, or timing. The study tracked reply rates across 45 of these groups separately from the content segments.
- Percentage stat segment
- The subset of emails whose body text includes a specific percentage figure as part of its pitch, such as a claimed lift or rate, the segment this study centers on.
- Largest group share
- The proportion of all emails in the dataset contributed by the single biggest sending group, used here to check that one campaign is not driving the whole result.
- Excluded rows
- Records removed from the dataset before analysis for reasons such as duplicates or malformed data. None were excluded in this study.
- Mitigation
- Any statistical adjustment applied after data collection to correct for imbalance or bias, such as reweighting. None were applied here, so the figures reflect the raw counts.
References
Sources this study reads against. Every link was fetched and confirmed reachable at publication.
- Cold Email Benchmark Report 2026 Instantly Establishes the industry-standard cold email reply rate baseline and defines reply rate as total replies over total emails sent.
- The Effect of Numbers on the Route to Persuasion Journal of Consumer Research (Oxford Academic) Foundational experimental study on how quantified claims interact with persuasion route, cited as the origin of the theoretical question our data speaks to empirically.
- When the 'Charm of Three' Fades: Mental Imagery Moderates the Impact of the Number of Ad Claims on Persuasion Journal of Consumer Psychology (Wiley) Shows that additional quantitative ad claims backfire under certain conditions via skepticism, an adjacent mechanism to our percentage-stat finding.
- How to Calculate Cold Email Reply Rates Instantly Summarizes Backlinko's 12-million-email outreach study reply rate and follow-up lift figures, used as an external benchmark for typical reply rates.
- Cold Email Response Rate (2026 Guide) Reachoutly Gives the standard counting rules (unique human replies, exclusion of autoresponders) that define what counts as a reply in this literature.
- Data-backed Tips to Write Subject Lines that Actually Work Klenty Finds subject lines with numbers substantially outperform those without, contradicting Belkins and illustrating dataset-to-dataset disagreement on this question.
- What are B2B Cold Email Response Rates? (2026 Study) Belkins Documents how a change in denominator (replies over opens vs. replies over sends) changed reported reply rates year over year, illustrating why measurement choices affect benchmark comparability.
- The asymmetrical force of persuasive knowledge across the positive-negative divide Frontiers in Psychology (PMC) Provides the skepticism/persuasion-knowledge mechanism, including the finding that claims beyond a small number increase skepticism rather than persuasion, used to explain why a percentage claim could plausibly reduce replies.
- B2B Cold Email Subject Lines and Engagement (2025 Study) Belkins Finds subject lines with numbers perform slightly worse than those without, the closest published result to our own finding that quantified content underperforms.
- I Analyzed 53M Cold Emails: 13 Stats That Matter in 2026 Saleshandy A large aggregated vendor dataset of cold emails and sequences, used as a comparison population for scale and methodology.
- 10 Data-Backed Ways to Boost Open Rates Yesware Reports that subject lines with numbers get a higher open rate than average, representing the prevailing industry claim our segment-level finding complicates.
- 15+ Email Subject Line Statistics [2023] Zippia Repeats the widely cited industry claim that numbered subject lines get much better open rates, used to characterize the strength of the prior belief we are testing against.
- Source effects in communication and persuasion research: A meta-analysis of effect size Journal of the Academy of Marketing Science (Springer) Meta-analytic estimate of how much credibility cues typically move persuasion outcomes, used as context for the magnitude of effects expected from a single copy element like a percentage.
Data
The complete row-level dataset is published open and ungated under CC BY 4.0. Every number in this study can be recomputed from it.
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
- Your category measured the same way
- Your own numbers, not a sample deck
- Fifteen minutes, no obligation
