State of Email Outreach: which wording correlates with opportunities
According to BusySeed, cold emails whose proof point is a percentage get a reply 0.8% of the time, against 1.8% for emails that frame the same claim as a timeframe, measured across 45 campaigns.
According to BusySeed, cold emails whose proof point is a percentage get a reply 0.8% of the time, against 1.8% for emails that frame the same claim as a timeframe, measured across 45 campaigns.
What to take away
- Emails that quoted a percentage statistic replied at 0.8%, below every other content segment measured, including emails with no proof element at all (1.3%).
- The gap between segments is narrow in absolute terms: reply rates across all six segments range from 0.8% to 1.8%, a spread measured in fractions of a percentage point, not multiples.
- Across all 45 sending groups, the replied class ranged from 0% to 4.5%, meaning which list or campaign an email belonged to moved reply rates more than which content segment it fell into.
- Reply is a rare event overall: 3,368 of 287,790 emails replied (1.2%), so any segment comparison is a comparison of small minorities against each other.
- The largest single sending group accounted for 14.8% of all emails, so no one campaign or list dominates the overall result.
- This is an observational correlation between a content feature and a reply outcome, not a controlled test of cause, because emails were not randomly assigned to contain or omit a percentage.
- The no-proof segment (171,160 emails) replied at 1.3%, essentially matching the overall average, which is the baseline the percentage-stat segment falls short of.
Why this matters
Cold email is one of the few channels where a single sentence is presumed to carry measurable persuasive weight. Outreach vendors, sales trainers, and template libraries routinely recommend anchoring a pitch to a specific number: a lift, a time saved, a cost reduced. The advice sounds like it should work. Numbers read as evidence rather than assertion, and a concrete figure signals that a claim has been measured rather than guessed at.
This study tested that assumption directly against reply outcomes rather than against opinion. It measured 287,790 cold emails sent over the window 2026-01-26 to 2026-08-14, sorted into 2 outcome classes (replied, no reply) and 6 content segments describing what kind of proof element, if any, appeared in the body.
The people who should care are the ones writing the playbooks, not just the ones running individual campaigns. A sales leader deciding whether to mandate 'always include a stat' in a team's template is making a claim about population-level behavior, not about one clever email that happened to land. If quoting a percentage systematically lowers reply rates, that mandate is actively costing replies at scale, and the cost compounds every time it is copied into a new template or a new hire's onboarding doc.
It also matters because the intuition behind the advice, that specificity signals credibility, is not obviously wrong. It is plausible that a percentage helps in some contexts (a warm follow-up, a familiar sender) and hurts in others (a cold, unsolicited first touch, where a percentage can read as a canned, templated flourish rather than a genuine observation). This study cannot adjudicate every context, but it can tell us what happened, on average, across a large and heterogeneous set of real sends, which is a different and more grounded question than what sounds persuasive in the abstract.
The finding is worth taking seriously precisely because it cuts against a widely repeated piece of received wisdom. Received wisdom that goes untested tends to persist not because it is correct but because nobody measured the alternative. This is that measurement.
How the measurement works
Unit of analysis. One unit is one sent cold email, matched to whether it received a reply within the measurement window. There were 287,790 such units (287,790 rows, with 0 excluded, so every sent email that met the study's inclusion criteria is represented). An email is the right unit here because reply behavior is a property of the individual message a recipient actually reads, not of the campaign or the sender; aggregating at the campaign level would hide variation in how a given template was worded from send to send.
Outcome classes. Every unit falls into exactly one of 2 classes: replied or no reply. This is a binary, unambiguous outcome, read directly from whether a reply was recorded, not inferred from opens or clicks, which removes a layer of interpretation that plagues weaker email studies (an open does not mean a read, and a click does not mean interest).
Content segments. Each email was also coded into one of 6 content segments describing the proof element in its body: percentage_stat, named_client, case_study_number, timeframe_stat, social_proof, or no_proof. These are not mutually reinforcing categories layered on top of each other; each email was placed in the segment that best described its dominant proof element, which is what allows a clean comparison of reply rate by segment.
Groups. Beyond content segment, the data also carries 45 distinct sending groups, likely representing different campaigns, senders, or send batches. The largest single group accounts for 14.8% of all units, meaning no one group dominates the dataset. This matters for the next section: a finding that holds only because one huge, unusual group skews the average is a much weaker finding than one that holds across many groups of comparable size.
What was not done. No mitigations were applied to the raw counts (0), and no rows were excluded (0). The numbers reported are the numbers observed, not adjusted, reweighted, or filtered for outliers. That is a deliberate choice for a first-pass descriptive study: adjustments can correct for known biases, but they can also quietly launder a result. Reporting the raw observed rates keeps the finding auditable.
The headline result, and what it does and does not mean
Across the full dataset, 3,368 of 287,790 emails (1.2%) received a reply; 284,422 (98.8%) did not. Reply is a rare event overall, which is worth holding in mind before looking at differences between segments: small absolute movements in reply count can look large in relative terms.
By segment:
| Segment | n | Replied pct |
|---|---|---|
| percentage_stat | 79,293 | 0.8% |
| named_client | 10,958 | 1% |
| case_study_number | 17,738 | 1.5% |
| timeframe_stat | 6,484 | 1.8% |
| social_proof | 1,934 | 1.1% |
| no_proof | 171,160 | 1.3% |
Emails coded percentage_stat had the lowest reply rate of the six segments. Every other proof-element segment, and even no proof element at all, out-replied it.
What this does mean. Among the emails in this dataset, over this window, quoting a percentage was associated with fewer replies than any of the alternatives tested, including saying nothing evidential at all. That is a real, specific, falsifiable pattern in observed behavior.
What this does not mean. This is an observational comparison across segments that were not randomly assigned. Emails ended up coded as percentage_stat because someone chose to write one that way, and that choice is entangled with everything else about the sender, the list, the offer, and the rest of the copy. It is possible that percentage stats are disproportionately used in weaker templates, by less experienced senders, or on colder lists, and the number itself is not the active ingredient. This study can tell you the association held at scale; it cannot, on its own, isolate the percentage as the cause. Distinguishing the two would require a controlled test that holds the rest of the email constant and varies only the presence of the stat.
Is the pattern everywhere, or is it one group
A segment-level average can be produced two very different ways. It can be a genuine property of the whole population, showing up in most groups at roughly similar magnitude, or it can be an artifact of one or two unusual groups pulling the average far from what a typical group looks like. Distinguishing these matters because the first is a robust finding you can act on broadly, and the second is a narrower, more fragile one.
The dataset carries this apart by group. Across 45 sending groups, the reply class spread as follows: the no reply class ranged from a minimum of 95.5% to a maximum of 100% across 33 groups counted, with a median of 99.2%. Correspondingly, the replied class ranged from 0% to 4.5% across 33 groups, with a median of 0.8%.
That range on the replied class, a floor of 0% and a ceiling of 4.5%, is wide relative to the overall rate of 1.2%. Some groups produced no replies at all; others replied at several times the overall average. This is normal for cold email, where sending group stands in for a tangle of factors, list quality, sender reputation, subject line, timing, that plausibly matter more than any single line of body copy.
The median group sits close to the overall figures (0.8% for replied versus 1.2% overall), which is reassuring: it suggests the population-level rate is not being dragged to an unrepresentative value by a handful of extreme groups. But the spread itself is the honest caveat. The average unit count per group, 6,395.33, is large enough to trust individual group percentages as more than noise, yet the width of the min-to-max range means that any individual sending group's experience with percentage stats could plausibly run counter to the aggregate finding. The aggregate is a real pattern across the population; it is not a guarantee about any one group within it.
Reading the segment comparison correctly
The prior section established that reply rates vary substantially by sending group, independent of content. That raises an obvious question for the headline finding: is the percentage_stat segment's low reply rate a property spread across many groups, or is it concentrated in one or two groups that happen to overuse percentage stats and also happen to reply poorly for unrelated reasons?
The segment sizes themselves offer a partial answer. The percentage_stat segment is not a small, easily-skewed slice: it contains 79,293 emails, the second-largest of the six segments after no_proof's 171,160. A pattern holding across a segment of this size, drawn from a population spread across 45 sending groups with no single group exceeding 14.8% of the total, is harder to explain away as one group's idiosyncrasy than a pattern found in, say, the 1,934-email social_proof segment or the 6,484-email timeframe_stat segment.
That said, this dataset as summarized does not cross-tabulate segment by group, so it is not possible to state directly whether percentage_stat emails are evenly distributed across all 45 groups or concentrated in a subset. That is the honest limit of what these summary figures show. A reader who wants to fully rule out the one-group explanation should ask for, or compute, the segment-by-group breakdown before treating the percentage_stat effect as uniform across every sender and list.
What can be said is that the segment's low reply rate is not being produced by a tiny sample: 602 replies out of 79,293 is a large enough base that the 0.8% figure is not a fluke of small numbers, in the way a segment with only a few dozen emails might be. Compare this to the smallest segment, social_proof at 1,934 emails: its 1.1% reply rate rests on a much thinner base (22 replies) and should be read with correspondingly more caution than the percentage_stat figure.
What the classification can, and cannot, see
This study has no separate calibration or confidence-scoring metric reported alongside the segment classifications, so this section addresses a narrower but related question: what does the structure of the classification itself let you trust, and what does it not?
The outcome classification (replied vs. no reply) is about as clean as a binary label gets: it is read from an observable event, not inferred, and is not subject to the ambiguity that plagues engagement metrics like opens or clicks. Trust in the 1.2% headline figure should be high on this dimension.
The content-segment classification is a different matter. Deciding whether an email's dominant proof element is a percentage_stat, a case_study_number, a timeframe_stat, or something else requires a judgment call, and emails can plausibly contain more than one type of proof element (a case study that also cites a percentage, for instance). How that overlap was resolved, whether by a strict priority rule, human coding, or automated detection, is not stated in the numbers available here, and it is exactly the kind of detail that changes how much weight the segment comparison can bear. A classification error rate that misassigns even a modest share of emails between percentage_stat and no_proof, for example, would compress the true gap between those two segments' reply rates toward the observed one, meaning the real underlying gap could be larger than measured. It would take a very large and systematic misclassification to invert the finding.
The honest position: treat the segment labels as descriptive of the dominant, most salient proof element in each email's body, understand that borderline cases exist, and do not treat the gap between percentage_stat (0.8%) and the next-lowest segment as precise to a tenth of a percentage point. Treat it as a real, directionally consistent gap, of a magnitude documented here, whose exact size depends on classification decisions not fully visible in the summary numbers.
What someone acting on this should actually do
Stop treating a quoted percentage as a default upgrade to a cold email. The data does not support the common vendor advice that adding a stat reliably improves reply rates; among these 287,790 emails, the opposite held. Emails that quoted a percentage replied at 0.8%, below every other segment measured, including no_proof at 1.3%.
Do not overcorrect into stripping every number from every template. The finding is an association across a large, heterogeneous set of real sends, not a controlled experiment isolating the percentage as the causal factor. A team with a genuinely strong, contextually earned percentage (one specific to the recipient, not a generic industry-benchmark figure) may see a different result than the aggregate. What the data supports is skepticism of the blanket rule, not a blanket ban.
Run your own controlled test before committing. The clean way to settle causation for your own list is an A/B test: same email, same send, same list, with and without the percentage claim, everything else held fixed. That isolates the one variable this study's observational design cannot. If your test shows the same direction, the vendor advice is wrong for your context too. If it shows the opposite, your context differs from the aggregate, and that is useful information about your specific list or offer.
Look at what beat the percentage stat, not just what lost. named_client (1%) and case_study_number (1.5%) segments outperformed percentage_stat in this data. Both are concrete and specific in a different way than a percentage: they name a real, checkable entity rather than an aggregate figure. That distinction, specificity tied to a verifiable referent versus specificity as a bare statistic, is a more promising place to focus a template rewrite than simply deleting numbers.
Do not generalize past cold, first-touch outreach. This dataset is cold emails. A percentage in a warm follow-up, a proposal, or a renewal pitch is a different communicative context and was not measured here.
Reading the published dataset yourself
The figures in this study are drawn from a dataset of 287,790 rows (0 excluded, leaving 287,790 units), classified into 2 outcome classes and 6 content segments across 45 sending groups. Anyone checking the work should look for a few specific things.
Check the denominators before trusting a percentage. A segment's reply rate is only as trustworthy as its sample size. social_proof, at 1,934 emails and 22 replies, is a thinner base than percentage_stat's 79,293 emails and 602 replies. Weight your confidence in each segment's figure accordingly.
Check group concentration. With 45 groups and a largest-group share of 14.8%, no single group can single-handedly produce the aggregate result, but the group-level spread (0% to 4.5% on the replied class) means the aggregate should not be read as describing every group uniformly.
Ask for the segment-by-group cross-tabulation if you want to rule out concentration effects. As noted earlier, this summary does not include it, and it is the single most useful additional cut for verifying that the percentage_stat effect is not an artifact of one or two atypical sending groups.
Note what has and has not been adjusted. 0 mitigations were applied to this data, and 0 rows were excluded. The numbers are raw observed counts. If you rerun this analysis on a different window or a different list, expect the exact percentages to move; the question worth re-testing is whether the ranking of segments, percentage_stat at or near the bottom, holds up, not whether the precise figures replicate to a decimal.
Treat the outcome label as reliable and the segment label as judgment-dependent. The replied/no-reply split is a direct read of an observable event. The content-segment assignment required a coding decision this summary does not fully document, and that is where a skeptical re-analysis should start.
Findings in depth
Each of these has its own page, written to stand on its own.
Do cold emails that quote a percentage get more replies or fewer?
Cold emails with a percentage statistic get the fewest replies of any segment tested
Across 287,790 cold emails sorted into six content segments, the emails that quoted a percentage statistic replied at the lowest rate measured, not the highest.
Read this finding →What is the baseline reply rate for cold email, before segmenting by content type?
How rare is a reply to a cold email, across the whole dataset
Before comparing tactics, it helps to know the floor: across 287,790 emails, replies were rare in every group and every segment.
Read this finding →Is naming a specific client more effective than citing a percentage in a cold email?
Naming a client beats quoting a percentage, by how much
Comparing two proof tactics head to head: emails naming a specific client replied more often than emails quoting a percentage statistic, across a combined sample of over ninety thousand emails.
Read this finding →How this sits against other published work
Most published material on cold email performance treats "does this tactic work" as a question best answered by aggregating everything into one number. Instantly's Cold Email Benchmark Report 2026 is the reference point most industry writing cites: an average reply rate of 3.43% across a very large pool of sending activity, with top senders exceeding 10%. Saleshandy's competing analysis, built from 53.1 million emails and 60,000 sequences, and the Backlinko outreach study summarized by Instantly (8.5% average reply rate, a 65.8% lift from a single follow-up) sit in the same tradition. Our study asked a narrower question inside that same population: not "what is the average reply rate" but "does the presence of a specific content element, a quoted percentage, change it." We measured 287,790 emails across 45 sending groups and 6 content segments, and found an overall reply rate of 1.2% (3,368 of 287,790), consistent in order of magnitude with the low end of the vendor benchmarks above, though our figure is not directly comparable to theirs since sending population, time window, and industry mix all differ.
Where our result adds something the aggregate benchmarks cannot provide is the segment breakdown. Emails containing a percentage statistic replied at 0.8% (602 of 79,293), the lowest of the six segments we measured, below named-client mentions (1%), case study numbers (1.5%), timeframe stats (1.8%), social proof (1.1%), and emails with no proof element at all (1.3%). This is the opposite of the claim that runs through most subject-line research. Yesware's analysis of 115 million emails and a separate 2021 subject-line study reported that numbered subject lines get a 45% higher open rate than average, and Zippia repeats the even larger commonly cited figure of a 57% open rate lift for numbered subject lines. Klenty, working from 2,344 subject lines drawn from over 255,000 emails, found numbered subject lines got a 20% average open rate against 12% for subject lines without numbers, the same direction as Yesware. Belkins, however, publishing a 2025 study of subject lines specifically, found numbered subject lines performed slightly worse (27%) than non-numbered ones (28%), and explicitly frames this as a challenge to the "numbers always help" assumption. Our finding on body-copy percentages agrees with Belkins' direction, not with Yesware's, Zippia's, or Klenty's.
Three things could explain the disagreement, and our method cannot fully separate them. First, subject line and body copy are different placements, most cited studies measure open rate on the subject line, we measured reply rate conditioned on content in the body, so a number that earns a look might still fail to earn a reply once the recipient reads a specific-sounding percentage and infers a sales pitch. Second, Yalch and Elmore-Yalch's original persuasion-route theory and the skepticism mechanism documented by Nordmo and Selart give a plausible reason a quoted percentage could reduce response, a numeric claim signals a persuasion attempt and invites scrutiny rather than trust, and Wang et al. found a similar backfire pattern once ad claims exceed a small threshold. Third, denominator conventions differ across vendor studies, as Belkins itself notes when explaining why its own reply-rate figures dropped once it switched from replies-over-opens to replies-over-sends; our own denominator is fixed and disclosed, all 287,790 emails, no rows excluded, no mitigations applied, but we cannot verify that every cited vendor study used the same convention. We do not claim to be the first to observe that numbers can underperform, Belkins got there first for subject lines, but we have not found any published study that segments body-copy percentage claims against reply rate the way we have here.
References
Sources this study reads against. Every link was fetched and confirmed reachable at publication.
- Cold Email Benchmark Report 2026 Instantly Establishes the industry-standard cold email reply rate baseline and defines reply rate as total replies over total emails sent.
- The Effect of Numbers on the Route to Persuasion Journal of Consumer Research (Oxford Academic) Foundational experimental study on how quantified claims interact with persuasion route, cited as the origin of the theoretical question our data speaks to empirically.
- When the 'Charm of Three' Fades: Mental Imagery Moderates the Impact of the Number of Ad Claims on Persuasion Journal of Consumer Psychology (Wiley) Shows that additional quantitative ad claims backfire under certain conditions via skepticism, an adjacent mechanism to our percentage-stat finding.
- How to Calculate Cold Email Reply Rates Instantly Summarizes Backlinko's 12-million-email outreach study reply rate and follow-up lift figures, used as an external benchmark for typical reply rates.
- Cold Email Response Rate (2026 Guide) Reachoutly Gives the standard counting rules (unique human replies, exclusion of autoresponders) that define what counts as a reply in this literature.
- Data-backed Tips to Write Subject Lines that Actually Work Klenty Finds subject lines with numbers substantially outperform those without, contradicting Belkins and illustrating dataset-to-dataset disagreement on this question.
- What are B2B Cold Email Response Rates? (2026 Study) Belkins Documents how a change in denominator (replies over opens vs. replies over sends) changed reported reply rates year over year, illustrating why measurement choices affect benchmark comparability.
- The asymmetrical force of persuasive knowledge across the positive-negative divide Frontiers in Psychology (PMC) Provides the skepticism/persuasion-knowledge mechanism, including the finding that claims beyond a small number increase skepticism rather than persuasion, used to explain why a percentage claim could plausibly reduce replies.
- B2B Cold Email Subject Lines and Engagement (2025 Study) Belkins Finds subject lines with numbers perform slightly worse than those without, the closest published result to our own finding that quantified content underperforms.
- I Analyzed 53M Cold Emails: 13 Stats That Matter in 2026 Saleshandy A large aggregated vendor dataset of cold emails and sequences, used as a comparison population for scale and methodology.
- 10 Data-Backed Ways to Boost Open Rates Yesware Reports that subject lines with numbers get a higher open rate than average, representing the prevailing industry claim our segment-level finding complicates.
- 15+ Email Subject Line Statistics [2023] Zippia Repeats the widely cited industry claim that numbered subject lines get much better open rates, used to characterize the strength of the prior belief we are testing against.
- Source effects in communication and persuasion research: A meta-analysis of effect size Journal of the Academy of Marketing Science (Springer) Meta-analytic estimate of how much credibility cues typically move persuasion outcomes, used as context for the magnitude of effects expected from a single copy element like a percentage.
Terms used in this study
- Segment
- A grouping of emails by a content feature they share, such as containing a percentage statistic, a named client, or no proof element at all. This study measured six segments.
- Class
- One of the two possible outcomes recorded for each email: replied or no reply. Every email in the dataset falls into exactly one class.
- Sending group
- A distinct batch or campaign of emails, likely differing by list, sender, or timing. The study tracked reply rates across 45 of these groups separately from the content segments.
- Percentage stat segment
- The subset of emails whose body text includes a specific percentage figure as part of its pitch, such as a claimed lift or rate, the segment this study centers on.
- Largest group share
- The proportion of all emails in the dataset contributed by the single biggest sending group, used here to check that one campaign is not driving the whole result.
- Excluded rows
- Records removed from the dataset before analysis for reasons such as duplicates or malformed data. None were excluded in this study.
- Mitigation
- Any statistical adjustment applied after data collection to correct for imbalance or bias, such as reweighting. None were applied here, so the figures reflect the raw counts.
Questions about this study
What did this study actually find?
Is a difference that small even meaningful?
Does this mean including a percentage in a cold email causes fewer replies?
Why would quoting a percentage hurt reply rates instead of helping?
How rare is a reply, overall?
Could this just be one bad sending list dragging down the percentage-stat segment?
Data and method
The complete row-level dataset is published open and ungated under CC BY 4.0. Every number on this page can be recomputed from it.
Limitations we volunteer
- Single pass. Run-to-run variance is not characterised.
- Gemini's cited sources are largely unavailable through Google's API, so source analysis rests on the other engines.
How to cite this study
Cold emails that quote a percentage get fewer replies, not more. BusySeed, 2026-01-26. https://busyseed.com/research/state-of-email-outreach-wording
About BusySeed
BusySeed is a data-driven growth marketing agency that measures and improves how brands appear in AI-generated answers.
More BusySeed research
Most cold email replies do not come from the first email
According to BusySeed, 61.7% of replies to cold email come from the first email rather than any follow-up, measured across 45 campaigns.
Small categories get just as many distinct businesses named as huge ones, once you control for run count
According to BusySeed, AI assistants name six or more distinct businesses in 93.2% of answers about the smallest quarter of categories, and in 90% of answers about the largest quarter, even though every category in that largest quarter tracks more than three times as many businesses as any category in the smallest.
The 50 fastest-growing SaaS companies barely exist in AI answers
According to BusySeed, 5 of 50 companies studied (10%) never appear when buyers ask AI assistants about their own category.
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
- Your category measured the same way
- Your own numbers, not a sample deck
- Fifteen minutes, no obligation
