🧪 Cold Email A/B Test Designer
Pick what you're testing, paste your control, and get your variants — plus a prospective sample-size estimate, the metric to track, and rules for when to call a winner. Works for subject lines and openers, or any two-to-three-arm test with a yes/no outcome where higher is better (trial vs demo CTA, trigger vs no-trigger, stage progression).
Set up the test
Test one variable at a time — otherwise you can't attribute the lift.
Higher is better — if your metric improves by decreasing (cancellation rate), enter its positive complement (retention rate).
Relative to the baseline: 30 means detecting 5% → 6.5%, not 5% → 35%.
Sample size — computed, not AI-estimated
4,578 per arm · 13,734 total · roughly 10–28 weeks at your volume
Two-sided two-proportion z-test: 5% → 6.5% at 95% family-wise confidence, 80% power per comparison. Three-arm sizing is Bonferroni-adjusted for the two control comparisons.
Your test plan
Your test design will appear here
Fill in the form and click Design the test
Saved test plans
Saved automatically on this device—no account required. Up to 20 per tool and 50 across all tools.
Your generated A/B test plans will appear here.
Other tools you might like
Frequently asked questions
How big does each test arm need to be?+
It depends on four things: your baseline outcome rate, the smallest lift you want to detect (adjustable in the tool, 30% relative by default), your confidence level (fixed at 95%), and statistical power (fixed at 80%). As a rough illustration, detecting a 30% relative lift on a ~5% baseline at those settings takes a few thousand subjects per arm — small lifts on low baselines need far more volume than people expect. The tool computes a prospective estimate from your inputs with a two-proportion z-test — deterministic maths, not an AI estimate — and shows it live as you type. Three-arm tests are sized with a Bonferroni-adjusted alpha, so the 95% confidence is family-wise across both control comparisons and each arm needs more subjects than in a two-arm test.
Can I size tests that aren't cold email tests?+
Yes. Any experiment with two or three arms and a yes/no outcome per subject works, provided higher is better for the metric: a trial CTA versus a demo CTA measured on stage progression, trigger-event emails versus a standard sequence measured on positive-reply rate, two landing-page headlines measured on signup rate. If your metric improves by decreasing, reframe it as its positive complement first. Pick "Something else" as the variable, name your outcome metric, enter its baseline, and choose 2 or 3 arms.
Should I test subject line and opener at the same time?+
No. Test one variable at a time. If you change both, you can't attribute the lift (or drop) to either. The tool enforces this — pick one variable and lock everything else.
What metric should I track?+
For subject-line tests: open rate (or reply rate if you don't track opens reliably — open tracking is increasingly unreliable with Apple Mail Privacy Protection). For opener tests: reply rate. For other experiments, use whatever binary outcome the test is meant to move — stage progression, meeting booked, signup — and enter it as the outcome metric so the plan and the sizing both use it. Pick a metric where higher is better: the tool sizes and judges tests for increases, so if yours improves by decreasing (a cancellation rate), reframe it as its positive complement (retention rate).
When can I call a winner?+
Only after every arm has hit the calculated sample size AND the winner's own comparison passes a significance test with the metric increasing — every plan declares higher-is-better, and a significant decrease means that arm lost. In a three-arm test, both planned variant-vs-control comparisons are evaluated at the Bonferroni-adjusted alpha of 0.025 (so the family-wise error rate is controlled at no more than 5%), but a variant wins on its own passing comparison — and B is never ranked against C without a planned direct comparison. The minimum detectable lift is a planning input for sizing the test, not the pass mark for the result. Calling winners early on small samples is the most common mistake in A/B testing.
Want more sales playbooks?
AI-assisted sales intelligence for closers, SDRs, and revenue leaders.
Read the latest articles →