Methodology

Most cold email replies do not come from the first email

Run 2025-12-01n = 1,857

Engines and exact versions

A study that does not name the model version it ran against is not reproducible, because the answer changes when the model does.

EngineModel version

Run window, UTC: [object Object].

Method

The unit. One unit is one reply. Each reply was traced back to the specific email in its sequence, the message it was sent in direct response to, and labeled with one of 4 classes: first email, follow-up 1, follow-up 2, or follow-up 3 or later. The dataset contains 1,857 such units, drawn from 1,857 rows with 0 rows excluded, spread across 45 outreach groups. A group is a distinct outreach campaign or sequence configuration; the average group contributed 41.27 replies, though as later sections show, group sizes and behavior varied.

Why attribute at the reply level, not the campaign level. A campaign-level count (how many total replies did this five-email sequence generate) cannot answer where in the sequence attention should go, because it collapses five distinct events into one number. Attributing each reply to the specific email that triggered it preserves the information a team actually needs: if someone is deciding whether to spend more time on email one or email three, they need to know how replies distribute across that specific choice, not just a campaign total.

What counts as a reply. Any inbound message a recipient sent back in the thread, regardless of sentiment, from an enthusiastic yes to an unsubscribe request, was captured as one unit and classified by sequence position. Sentiment is tracked separately as a segment (interested, neutral, not interested, unlabelled, unsubscribe or hostile) and covered in a later section. This separation matters: a study that only counted positive replies would answer a different, narrower question, since it is possible that positive replies cluster differently than hostile ones across the sequence.

Classification and its limits. No mitigations were applied to the data (0), meaning the class assignments are taken as recorded, without adjustment for possible mislabeling. The measurement window ran from 2025-12-01 to 2026-08-14. What this design cannot see: it cannot tell you whether a follow-up reply would have arrived anyway without the follow-up being sent, nor whether a different subject line on follow-up two would have pulled more replies. It measures where replies landed in the sequences that were actually run, not what would happen under a different sequence design.

Limitations we volunteer

Written by us, before anyone else found them.

  • Single pass. Run-to-run variance is not characterised.
  • Gemini's cited sources are largely unavailable through Google's API, so source analysis rests on the other engines.

Terms used in this study

First email
The initial message in a cold outreach sequence, sent before any follow-up.
Follow-up 1, 2, 3 or later
Subsequent messages sent to the same recipient after the first email, numbered in the order they were sent. Follow-up 3 or later groups together all replies to the third follow-up and any sent after it.
Class
In this study, one of the four positions in the sequence, first email, follow-up 1, follow-up 2, or follow-up 3 or later, that a reply is attributed to based on which message it was a response to.
Segment
A grouping of replies by the sentiment of the recipient's response, such as interested, neutral, not interested, unsubscribe or hostile, or unlabelled when sentiment could not be determined.
Group
One outreach campaign or batch within the dataset. The study covers 45 groups, and results are also checked group by group to see whether the overall pattern holds consistently or is driven by a few groups.
Largest group share
The proportion of all replies in the dataset contributed by the single biggest group, used to check that the headline result is not driven by one unusually large campaign.
Mitigations applied
Any adjustments made to the raw data to correct for known biases or errors before analysis. None were applied in this study.

References

Sources this study reads against. Every link was fetched and confirmed reachable at publication.

  1. Cold Email Response Rates: B2B Benchmarks Instantly, 2026 Cross-references the Backlinko and Belkins figures alongside Instantly's own benchmark, useful for comparing independent reply-rate baselines.
  2. Cold Email Benchmark Report 2026: Reply Rates, Deliverability and Trends Instantly, 2026 Source of the widely repeated claim that 58% of replies come from the first email in a sequence, based on Instantly's own platform data.
  3. How to Calculate Cold Email Reply Rates Instantly, 2026 Defines reply rate as human replies divided by delivered emails and describes segmenting replies by sequence step, the same attribution logic our study depends on.
  4. We Analyzed 12 Million Outreach Emails. Here's What We Learned Backlinko, n.d. Foundational independent study establishing an overall cold outreach response rate, cited by nearly every derivative benchmark including Instantly's.
  5. What are B2B Cold Email Response Rates? (2026 Study) Belkins, 2026 Documents a year-over-year methodology change that lowered Belkins' own reported reply rates, relevant to assessing whether cross-study differences reflect measurement artifacts.
  6. Sales Follow-Up Statistics in B2B (2026 Study) Belkins, 2026 Reports that follow-ups collectively produce the majority of replies even though the first email has the highest per-step rate, the main disagreement in framing with our headline result.
  7. I Analyzed 53M Cold Emails: 13 Stats That Matter in 2026 Saleshandy, 2026 Independent large-sample study concluding that most positive replies come from follow-ups rather than the opener, directly contradicting the Instantly framing.
  8. Cold Email Statistics Based on Sending Over 20M Cold Emails Woodpecker, 2026 Widely cited industry baseline reporting platform-wide reply rate decline and recommended follow-up timing, representing the 'reader's prior' baseline.

Data

The complete row-level dataset is published open and ungated under CC BY 4.0. Every number in this study can be recomputed from it.

Free strategy session

Want to know how AI answers describe you?

We run the same measurement on your category. Fifteen minutes with founder Omar Jenblat, your own numbers, no deck.

Omar Jenblat, Founder & CEO of BusySeed
Omar JenblatFounder & CEO, BusySeed
  • Your category measured the same way
  • Your own numbers, not a sample deck
  • Fifteen minutes, no obligation

First, who are we meeting?

Three fields, then pick your time. We read up on you before the call so we open with something useful.

No sales sequence. If you never pick a time, we leave it there.