Why your cold email A/B test is lying
Your test did not go wrong because the sample was small. It went wrong because of what you and your sending tool did to a small sample once the replies started arriving: the looks, the extra variants, the split that was never a coin flip, the denominator that moved underneath it. Six mechanisms, and a check for each one.
By Kshitij Maheshwari, co-founder · Updated August 2026 · 12 min read
The post-mortem, in six checks
Run these against a result you already believe. If the test is not designed yet, start with our guide to A/B testing in outbound.
| The check | A bad answer | What it costs |
|---|---|---|
| How many times did you look before calling it? | "I checked most mornings." | Daily looks for four weeks turn a 5% test into a 25.6% one. |
| How many comparisons were live at once? | "Four subject lines, then I read it by segment." | Six comparisons, and a 26.5% chance one looks like a winner. |
| Did the split land where you set it? | "Roughly even, I think." | At 1,000 contacts a 52/48 split passes every check you can run. |
| Does the tool carry a variant through the sequence? | "Version A is the whole sequence." | Two platforms document that it does not. |
| What sat on the bottom of the fraction? | "Reply rate. It is in the dashboard." | Bounces, opt-outs and mid-flight edits never land evenly. |
| What could the test have seen if the gain were real? | "It was significant, so it was real." | At 500 per arm on 3%, the chance of seeing it runs 5% to 24%. |
Sources: Evan Miller (2010) for what stopping early costs, and Instantly's and Apollo's own help pages (checked August 2026) for what the two platforms do. Every percentage in the table is ours, computed from the inputs printed beside it.
Looking at it every day is the whole problem
Opening the campaign on a Tuesday morning is the most normal thing a founder does.
The significance figure on your dashboard is only true if you fixed the size of the test before the send. Evan Miller named that assumption in 2010: the calculation "makes a critical assumption that you have probably violated without even realizing it: that the sample size was fixed in advance."
Set the stopping line before you send
One reading, at a point fixed before the send. That is the 5.0% you signed up for, and the row proves the simulation behaves.
Drop the Monday morning review
Four weeks of it, stopping the moment the test reads significant, lands at 12.9%. The dashboard's 5% is already understated by more than double.
Stop opening it every few days
Ten looks takes it to 19.4%, so one dead-even test in five hands you a winner. Miller's correction table agrees: at ten peeks, 1% is really 5%.
Stop deciding on the morning refresh
One look a day for four weeks, and 25.6% of tests name a winner where the two arms are identical by construction. The habit, priced.
report a significant winner, when the arms are identical and someone checks every morning for four weeks and stops at the first significant look.
Two arms of 1,000 contacts, both running at a 3% reply rate, with no real difference between them.
A low reply rate does not soften this, and neither does volume. Miller's worst case, at a 50% reply rate, lands at 26.1%. The work behind Optimizely's own statistics engine found that even at 10,000 contacts a side, peeking runs the false-winner rate to roughly five times what the dashboard prints.
Refreshing the dashboard before coffee is not carelessness, it is caring, and I still do it. What I stopped doing is deciding on what I saw. We write the stop line before the send, and where a number is too thin to carry a claim, we leave the claim off.
Four variants is six comparisons
Five percent is what one pre-committed comparison costs, and nobody makes one comparison.
| What you ran | Comparisons | Chance one looks significant when nothing is happening |
|---|---|---|
| Two variants, one metric, read once | 1 | 5.0% |
| Three variants | 3 | 14.3% |
| Four variants | 6 | 26.5% |
| Three metrics read across four segments | 12 | 46.0% |
That column is one minus 0.95 raised to the number of comparisons, computed by us, and four variants make six pairs rather than four. It holds only where comparisons are independent. Repeated looks at one running test are not, so ten looks is about 19%, never 40%.
So the move is two variants and one metric, both named before the send. If you also want to read the result by segment, that is a second test with its own list, not a second look at this one.
The split you configured is not the split you got
Experimentation teams void a run over an unexpected split. At your volume you cannot detect one.
An off split usually voids the result
The experimentation teams at four software companies wrote this down together in 2019: a split that lands off target "in most cases completely invalidates experiment results". They throw the run away.
The check cannot see a test your size
Run their own check on 1,000 contacts and it only fires past about 53 against 47. A 52/48 split, plenty to bend a result at your volume, sails straight through it.
Build the split, do not test it
Shuffle the file before splitting it, split inside each mailbox and each send day rather than across them, and count both arms before the first send.
The split was rarely a coin flip to begin with: lists arrive in scrape order, one arm lands on a newer domain, the halves go out into different deliverability weather. HubSpot documents its own version of this: sends "may not be 50/50 (especially with smaller lists)".
Two tools re-randomize at every step
If you tested two sequences rather than two emails, read your platform's A/B page first.
"Sequence A beat sequence B by two points, so we are rolling out sequence A."
- ✕Assumes one lead sees one version end to end
- ✕Assumes step two belongs to the same arm as step one
- ✕Nothing in the dashboard says otherwise
"A lead who receives Variant A in Step 1 could receive any variant (A, B, C, etc.) in Step 2. This is by design."
- ✓Instantly, on its own A/Z testing page, checked August 2026
- ✓Apollo documents the same behavior, updated 22 July 2026
- ✓So the unit you tested was the step, never the sequence
Neither hides it: both publish the behavior on their own help pages. What no dashboard does is size the test, and 100 or 200 delivered messages is a habit rather than a sample size, so compute the number before you send.
Instantly publishes a second fact worth knowing: it balances variant usage across a campaign's whole life, so add a variant mid-campaign and "only the new variants will be sent until their usage catches up".
Sources: Instantly Help Center, A/Z Testing (undated, checked August 2026); Apollo Knowledge Base (22 July 2026).
The most useful three minutes in a testing week go on your platform's own A/B help page. Every founder can do it and almost nobody has. The rule it prints is rarely the rule you assumed the dashboard was following.
Want outbound run by two people who write the decision down first?
Book a Fit CheckYour denominator moved while you watched
A rate has two moving parts and only one is replies. The question is not which denominator to pick, it is that yours did not hold still.
Quote the rate the dashboard hands you
Variant A 4.2% / Variant B 2.8% / Winner: A
- ✕No record of what sat on the bottom
- ✕Bounces left the arms different sizes
- ✕One thread counted twice on one arm only
Write the whole fraction on the day you call it
A: 21 people replied of 500 enrolled on 3 Aug, 12 bounced, 2 opted out, each person counted once
- ✓Top and bottom fixed on one date
- ✓The same counting rule on both arms
- ✓A mid-flight edit shows up instead of hiding
An illustrative walkthrough of the mechanism, not a specific client result. We report real numbers only when they are real. The example numbers are invented.
A test this small mostly cannot see anything
Before asking whether the winner is real, ask what the test could have found if it were.
7 replies against 12, at 100 sends each
A result shaped like that gets called a clear winner and rebuilt around. It is not one. At 100 sends a side, a gap that size turns up regularly when the two emails are identical.
500 a side, at a 3% reply rate
Every figure below assumes two arms of 500 contacts and a 3.0% reply rate on the control. The arithmetic is ours, and it is the ordinary version of it.
It misses almost any gain worth chasing
Against a real 5% improvement, the chance it spots the gain is 5.2%. At 20% it is 8.3%, at 50% it is 23.9%. Under a 15% improvement it never clears 10%.
Read what the test could see, not what it saw
Numbers like these swing on a guess nobody can check in advance. At a real 10% improvement, 21.9% of the winners that clear the bar are the worse arm. At 5% it is 34.5%.
That chance of spotting a real gain is the number to lead with, because it stays low across every improvement anyone would plausibly claim. Two things go wrong underneath it, and both of them flatter you.
A test this thin can hand you a real winner dressed as the loser. And any winner that does clear the bar looks bigger than it is.
Five hundred a side is nowhere near enough to keep either of those off you. Separating a 3% reply rate from a 4.5% one takes about 2,500 contacts per variant. The volume table behind that price is on our guide to outbound market learning. That is the argument on that page. This is what follows from it.
The winner that shrinks when you run it again
Everything above is diagnosis. This is the one thing here you can test in a week.
Run the winner against the loser once more, on a fresh slice, re-randomized, and expect the gap to shrink. A margin that cleared the bar from that position is mostly the luck that got it there.
Re-randomize rather than reusing the same two halves. Reuse them and whatever the first test did to those contacts, who opened it, who was already mid-thread with you, rides into the second one and reads back as a result.
The same discipline applied to audiences rather than wording is our play on ICP slice experiments, which confirms a slice with a second batch.
Run one A/A test before you trust the next one
Spend one campaign finding out what your tool reports when there is nothing to report.
-
1
Send the identical email to both halves
Same copy, same subject line, same mailboxes, same send days. Label them A and B exactly as you would for a real test.
-
2
Run it the way you run a real one
Same volume, same duration, same number of morning check-ins. You are testing your procedure as much as your tool.
-
3
Read the gap as your noise floor
Whatever it reports came from your split, your mailboxes or your send days, because the copy was identical. Future winners have to beat it.
Ron Kohavi said it plainly back in 2012: "The A/A test has been our most useful tool in identifying issues in practical systems." One campaign is the whole price.
Five ways a result lies to you
The shapes those mechanisms take by the time they reach you, four of them dressed as good news.
It lost overall but won with ops leaders, so the test "worked". A segment chosen after reading the data is another comparison.
You had reply rate, positive reply rate and meetings booked. Leading with the one that moved costs the same as running three tests.
Worked out after the result is in, that number restates the result in different clothes. It cannot tell you whether the test was ever able to see what you are claiming.
Triple-digit lifts get published off a handful of replies on each side. The thinner the test, the more of the gap is the luck that produced it, and the less of it survives a re-run.
Auto-optimize will switch an arm off for you
Instantly documents that the feature "automatically analyzes variant performance and deactivates lower-performing versions, based on a defined winning metric", and open rate is one of the choices. A deactivated arm stops collecting replies, so the call cannot be revisited.
What to do with all of that
A result is not a measurement until you can count the decisions that produced it.
- 1 Fix the sample size before the send, then read the result once.
- 2 Two variants, one metric. Every extra comparison buys another false winner.
- 3 Count both arms yourself. The split check is blind at your volume.
- 4 Re-run the winner on a fresh split before you rebuild around it.
Questions founders ask
How many emails do you need per variant for a cold email A/B test?
Is it bad to stop a test early when one version is clearly winning?
How many variants should I run at once?
Should I trust the winner my sending tool picked?
What is an A/A test and why would I run one?
How do I tell a real winner from a lucky one after the fact?
Co-founder of Real Good GTM. He has been the first business hire and Chief of Staff at seed-stage B2B startups, building outbound pipeline before any playbook existed. He wrote this one because the argument he has most often with founders is about a winner that was never there, and the arithmetic settles it faster than opinion does.
Connect on LinkedInBefore you run the next one
The design side of the problem, the arithmetic behind the volumes, and the play that re-runs its own results.
Cold email A/B testing
What to randomize, what to freeze, and the one decision rule to write down before the send.
Read the guideOutbound market learning
What each send volume actually buys you, with the interval table and the denominator ledger.
Read the guideICP slice experiments
Testing audiences instead of adjectives, with a confirmation batch built into the motion.
See the playWant outbound where the numbers survive a second look?
Book a fit check. We'll look at what your current numbers can and cannot support, set the one decision each campaign is allowed to make, and tell you straight if outbound is not the right motion for you yet.
Book a Fit CheckNo hard sell. No fake numbers. Real good work speaks for itself.