Most cold email replies come from the first email, not the follow-ups
Across 1,857 replies, the first email in a sequence generated most of the responses, and each additional follow-up produced sharply fewer.
61.7%
If you send a cold email sequence of four messages, where do the replies actually come from? This study measured that directly by tagging every reply in a dataset of 1,857 replies, collected across 45 outreach groups between 2025-12-01 and 2026-08-14, with the position in the sequence that produced it: first email, follow-up 1, follow-up 2, or follow-up 3 or later. No rows were dropped and no mitigations were applied to the counts, so these are the raw shares.
The answer is lopsided. 1,145 of 1,857 replies, or 61.7%, came from the first email alone. Follow-up 1 produced 388 replies (20.9%), follow-up 2 produced 233 (12.5%), and everything from follow-up 3 onward produced only 91 replies (4.9%). Each step down the sequence roughly halves the share of replies it generates, and by the third follow-up the contribution is small enough that in some outreach groups it did not appear at all (see the group-level spread below).
Why would this be true? A cold email sequence is, mechanically, a series of prompts to the same inbox, and the population of recipients willing to respond shrinks with every message that doesn't land. The first email reaches everyone in the list at full novelty: it is the first time the recipient has seen the sender's name, the offer, and the ask. Anyone who was going to respond because the pitch was relevant, well-timed, or simply worth a two-line reply tends to do so immediately, because there is no cost to waiting and no reason to defer a reply you've already decided to send. By the time follow-up 1 arrives, the remaining pool is disproportionately people who did not find the first email compelling enough to act on. Follow-up 2 draws from a pool that has now ignored two messages. The math of diminishing returns is not a claim about follow-ups being poorly written; it is a structural consequence of who is left in the recipient pool at each stage.
What would make this false? If follow-up emails were, on average, better copy than the first email, or if they reached a systematically different and more receptive audience (for example, a follow-up sent after a trigger event, or one routed only to recipients who opened but didn't reply), you would expect the decay to flatten or reverse. It doesn't here. The decline is monotonic across all four classes, and it holds up when the replies are split by sentiment segment: even among unsubscribe-or-hostile replies, the first email accounts for the largest share. That consistency across segments is evidence the pattern is about sequence position, not about the content of any particular first email or follow-up in this dataset.
What this does not tell you. This measures where replies land, not whether follow-ups are worth sending. A follow-up that generates only 4.9% of total replies can still be the deciding factor for the recipients it does reach, and this dataset cannot tell you whether those particular replies would have happened anyway, later, without the follow-up. It also cannot tell you the reply rate per email sent at each stage, only the share of replies attributable to each stage, since fewer recipients remain in the sequence by the time follow-up 3 rolls around. A shrinking pool naturally produces a shrinking numerator even if the per-recipient reply rate held steady.
What a company does differently on Monday. If the first email is where most of the reply volume originates, then the marginal hour spent sharpening the first email's subject line, opening sentence, and specific ask is likely worth more than the marginal hour spent polishing follow-up 3. That doesn't mean cutting follow-ups; the decay curve here still leaves a meaningful tail of replies in follow-up 1 and follow-up 2 combined (20.9% and 12.5%). It means treating the first email as the primary lever and follow-ups as a smaller, secondary one, and measuring them separately rather than judging a sequence's quality only by its aggregate reply rate.
Want to know how AI answers describe you?
We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.
- Your category measured the same way
- Your own numbers, not a sample deck
- Fifteen minutes, no obligation
