Questions about this study

Most cold email replies do not come from the first email

10 questions answeredMethod challenges included
What did this study actually find?
Across 1,857 cold email replies collected between 2025-12-01 and 2026-08-14, 1,145 of them, 61.7% of the total, came from the first email in the sequence rather than from any follow-up. Follow-up 1 accounted for 20.9%, follow-up 2 for 12.5%, and follow-up 3 or later for 4.9%. The data spans 45 outreach groups averaging 41.27 replies each, with no rows excluded and no mitigations applied. In short, most of the reply volume a sequence will ever generate arrives before the first follow-up is even sent.
Does this mean follow-ups are a waste of time?
No, and the study does not claim that. It only measures where replies land within a sequence, not why, and not what would have happened if follow-ups had been skipped entirely. Follow-up 1 and follow-up 2 still bring in 20.9% and 12.5% of replies respectively, which is not nothing. It's possible some of those replies would never have arrived without the earlier email as groundwork. Distinguishing 'follow-ups add incremental replies' from 'follow-ups just catch people who were always going to reply eventually' would require a design that holds sequence length constant and compares against a no-follow-up control, which this study is not.
Is the pattern different for people who actually seemed interested, versus people who ignored or rejected the outreach?
The pattern holds across all sentiment segments, but with some variation. Among the 123 replies labeled interested, 52.8% came from the first email. Among the 391 not-interested replies, it was 55.2%. The most front-loaded segment by far was unsubscribe or hostile replies: 87.1% of just 31 such replies arrived after the first email alone, meaning people who wanted to opt out or push back rarely waited for a follow-up to say so.
Why would the first email get so many more replies than the later ones?
The study doesn't test mechanism directly, but a few explanations are plausible. Recipients who are going to respond to cold outreach at all may do so at the first opportunity, before inbox fatigue or annoyance sets in. It's also likely that each successive follow-up is sent to a shrinking pool, people who already replied or unsubscribed are removed from later stages, so there are structurally fewer recipients left to generate a follow-up-3 reply. The data can't separate 'people reply faster to first contact' from 'there are just fewer people left by follow-up three', and those two explanations call for different fixes.
How much does the first-email share vary across different outreach groups?
Considerably. Across the 14 groups measured, the share of replies attributed to the first email ranged from a low of 46.6% to a high of 89.5%, with a median of 69.4%. That's a wide spread, which means the headline figure of 61.7% is an average, not a guarantee for any individual campaign. Sequence design, audience, and subject matter likely all move this number, though the study doesn't isolate which factor matters most.
Could one huge campaign be driving this whole result?
That doesn't appear to be the case. The largest single outreach group contributed only 27.9% of all 1,857 replies in the dataset, and the average group contributed 41.27 replies out of 45 groups total. Because no group dominates, the front-loaded pattern isn't an artifact of one unusually large or unusual campaign skewing the aggregate. It's a pattern that shows up broadly, even if its exact strength varies group to group, as the wide spread in first-email share across groups shows.
What counts as a 'reply' here, and how were replies assigned to a class?
A reply is any response a recipient sent back to a cold email in a sequence, and each reply was assigned to exactly one of 4 classes based on which email in the sequence it was responding to: first email, follow-up 1, follow-up 2, or follow-up 3 or later. The study reports 1,857 total rows with 0 excluded, meaning every reply in the collection window was classified and none were dropped for data quality or other reasons. The classification itself, which email a reply 'belongs to', is a matter of matching the reply to the message it responded to.
What should a sales or growth team actually change based on this?
The finding argues for putting more effort into the first email, since it is responsible for the majority of replies a sequence generates, 61.7% of the 1,857 in this dataset. Concretely, that could mean spending more time on the opening subject line and first message body relative to later touches, or testing whether a stronger first email raises the total reply rate more than adding another follow-up would. What it does not tell you is whether shortening sequences altogether would help or hurt, since the study only shows where replies land within existing sequences, not what would happen if follow-ups were removed.
Does 'more replies from the first email' mean the first email is more persuasive?
Not necessarily. It means more replies are attributed to it, which could reflect persuasiveness, but could equally reflect timing, attention, or list attrition. A recipient who was always going to respond might simply do so at first contact rather than waiting. Also, by the time follow-up 3 goes out, the pool of remaining recipients has been thinned by earlier replies and unsubscribes, so there are fewer people left who could generate a follow-up-3 reply regardless of how persuasive that email is. Testing persuasiveness specifically would require comparing reply rates per email holding recipient pool size constant, which this study does not do.
Is this result about reply volume, or about the quality of the leads?
It's primarily about volume and timing, not quality, though the segment breakdown gives a partial view into quality. Interested replies show a somewhat lower first-email share, 52.8% of 123, than the overall 61.7%, meaning a slightly larger fraction of interested responses arrive at later stages compared to the full dataset. That's a real difference worth noting, but it is a modest one relative to sample size, and the study doesn't measure downstream outcomes like meetings booked or deals closed, so it cannot speak to lead quality beyond the interested and not-interested labels themselves.
Free strategy session

Want to know how AI answers describe you?

We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.

Omar Jenblat, Founder & CEO of BusySeed
Omar JenblatFounder & CEO, BusySeed
  • Your category measured the same way
  • Your own numbers, not a sample deck
  • Fifteen minutes, no obligation

First, who are we meeting?

Three fields, then pick your time. We read up on you before the call so we open with something useful.

No sales sequence. If you never pick a time, we leave it there.