Sales Email Subject Lines: What Research Shows
What Boomerang, Salesloft, Apple, and deliverability documentation support about subject-line length, reply metrics, and properly powered tests.
Review note: Checked the Boomerang, Salesloft, and Lavender evidence, open-tracking limitations, and the statistical assumptions behind the proposed subject-line tests.
What the published evidence actually covers
Public subject-line research supports a narrower conclusion than most best-practice lists suggest: shorter subject lines are reasonable candidates to test, but no public dataset establishes a universal winning formula for cold email.
Boomerang analysed more than 40 million emails for which its customers had asked to be reminded if no response arrived. Subject lines containing three or four words received the most replies in that dataset, and response rates declined gradually as more words were added.
That is useful observational evidence, with important limits. Boomerang published the analysis in 2016, covered email generally rather than cold B2B outreach, and did not present the subject-line result as a randomised experiment. The page also does not disclose enough segmentation or modelling detail to show that length caused the difference.
Salesloft's subject-line analysis, first published in 2019 and updated in 2023, says it drew on hundreds of millions of interactions in its sales-engagement platform. Its summary recommends one to four words and reports weaker reply-rate associations once subject lines reach six words. This is closer to the B2B sales context, but it remains vendor-reported observational data rather than a controlled comparison.
Lavender's current guidance also recommends very short, descriptive subject lines. The page recommends one to three words, with two described as optimal, but it does not publish a subject-line sample definition or statistical methodology. It also cites Salesloft for some findings. Treat it as vendor guidance, not independent replication.
The defensible shared finding is that brevity deserves a place in a test plan. It does not prove that a short subject line will outperform a longer, clearer one for every audience.
Where the popular patterns conflict
The public sources do not support presenting first names, questions, lowercase fragments, and numbers as four patterns that consistently win. In fact, the vendor datasets sometimes point the other way.
- First-name-only subjects: Salesloft associated the recipient's first name in the subject line with a reply rate below its sample average. A first-name-only line can also resemble an internal message, so it should not be the default treatment.
- Questions and numbers: Lavender advises against both in its current guide. Salesloft also reports a negative association for numbers. Those findings do not prove that every relevant number or question performs badly, but they directly contradict a blanket claim that the patterns win.
- Lowercase fragments: Lavender recommends title case, while Salesloft reports all-caps subject lines outperforming all-lowercase ones, although Salesloft explicitly advises against using all caps. The disagreement is a reason to test casing, not to declare lowercase the winner.
- Verified trigger details: A real product launch, hiring change, filing, or public statement can make the subject accurately preview the email. Neither linked dataset proves that adding such a detail increases replies. Treat the idea as a testable editorial recommendation.
Avoid deceptive Re: or Fwd: prefixes when there is no earlier thread. That is an editorial standard about truthful presentation, not a claimed performance finding.
The same distinction applies to words such as "free" and "guarantee." Salesloft reports reply-rate associations for promotional language, but that does not establish that a single word caused a spam filter to block the message. Delivery depends on factors including authentication, sender reputation, complaint history, recipient behaviour, and message content. Measure delivery and complaints rather than relying on a list of supposed trigger words.
Why open rate needs a warning label
Open tracking usually depends on a remote image request. Apple explains that Mail Privacy Protection can download remote content in the background whether or not the recipient engages with the message. That prevents the sender from knowing whether an apparent Apple Mail open represents a human view.
Security infrastructure adds another source of noise. Twilio SendGrid documents that spam filters, recipient bots, and prefetching systems can generate non-human opens or clicks that appear in engagement reporting.
Open rate can still help diagnose a large change within a stable sending environment, particularly when a platform identifies machine activity. It should not be treated as a clean human-engagement measure or the sole criterion for choosing a subject line.
For prospecting tests, use delivered messages as the denominator and make reply rate or positive reply rate the primary outcome. Track bounces, complaints, and opt-outs as guardrails. Define in advance how out-of-office replies, automated messages, and ambiguous responses will be classified.
Replace the 400-send rule with a sample-size plan
There is no defensible universal rule that 400 sends per variant can detect a two-percentage-point improvement. Required sample size depends on the baseline rate, the smallest effect worth detecting, the number of variants, the desired false-positive rate, and statistical power.
For example, a two-sided two-proportion planning approximation with 95% confidence and 80% power requires about 2,213 delivered messages per variant to distinguish a 5% reply rate from a 7% reply rate. At a 3% baseline and a 5% alternative, the same approximation requires about 1,506 per variant. These are illustrative planning calculations, not guarantees about a completed experiment.
Use the SalesTap A/B Test Designer to state the baseline, minimum detectable lift, variants, and planned sample before sending. If the required sample is unrealistic, either accept that the result will be directional or design a test around a larger commercially meaningful effect. Do not keep checking the result and stop as soon as one variant moves ahead.
The measurement window also needs to be chosen before the test. Seven days may be reasonable for one sequence and too short for another. Use a window that captures the normal reply lag for the audience, then apply it consistently to every arm.
A cleaner subject-line experiment
Start with one claim the public evidence can support testing: whether a short, descriptive subject line performs differently from the current control.
- Keep sender, audience, offer, email body, timing, and follow-up sequence as similar as practical.
- Randomly assign comparable recipients to the control and variant during the same period.
- Change only the subject line. A useful variant is one to four words and accurately describes the email topic.
- Pre-register the sample size, reply window, exclusions, primary metric, and guardrails.
- Check delivered volume and bounce rates by arm before interpreting replies.
- Report the observed rates and uncertainty. Do not describe a small numerical lead as a winner unless the planned analysis supports it.
If the existing control is already short, test one different attribute at a time. Casing, a verified trigger detail, or a question can each be a separate treatment. Combining all three in one variant would show whether the package differs, not which element mattered.
The takeaway
- Test brevity rather than treating it as law. Boomerang, Salesloft, and Lavender all favour short subject lines, but their evidence and populations differ.
- Do not call the four popular patterns proven winners. The current vendor guidance conflicts on names, questions, numbers, punctuation, and casing.
- Use replies with safety guardrails. Apple privacy features and security scanners make opens an imperfect proxy for human attention.
- Plan the sample from the baseline and target effect. A fixed 400 sends per variant is often too small for the difference a team says it wants to detect.
- Keep the subject truthful. A clear description of the email topic is a safer starting treatment than manufactured familiarity or a deceptive thread prefix.
Put this into practice
Use our free AI tools to apply these tactics immediately.
Explore free sales tools ↗Keep reading
Cold Call Openers That Get Buyers Talking
The seven seconds after a prospect picks up determine everything — here's how elite SDRs open cold calls and keep the conversation alive.
LinkedIn Prospecting Without Annoying Buyers
Stop blasting pitches into DMs. Here's the trigger-based, value-first LinkedIn framework that earns responses from serious B2B buyers.
6 Cold Email Templates for MSPs (With Triggers)
Six MSP cold email templates built around verifiable hiring, insurance, acquisition, compliance, and security triggers, with a practical testing plan.