Skip to content
From the blog

Why your cold email A/B test is lying

Your test did not go wrong because the sample was small. It went wrong because of what you and your sending tool did to a small sample once the replies started arriving: the looks, the extra variants, the split that was never a coin flip, the denominator that moved underneath it. Six mechanisms, and a check for each one.

By Kshitij Maheshwari, co-founder · Updated August 2026 · 12 min read

The post-mortem

The post-mortem, in six checks

Run these against a result you already believe. If the test is not designed yet, start with our guide to A/B testing in outbound.

The check A bad answer What it costs
How many times did you look before calling it? "I checked most mornings." Daily looks for four weeks turn a 5% test into a 25.6% one.
How many comparisons were live at once? "Four subject lines, then I read it by segment." Six comparisons, and a 26.5% chance one looks like a winner.
Did the split land where you set it? "Roughly even, I think." At 1,000 contacts a 52/48 split passes every check you can run.
Does the tool carry a variant through the sequence? "Version A is the whole sequence." Two platforms document that it does not.
What sat on the bottom of the fraction? "Reply rate. It is in the dashboard." Bounces, opt-outs and mid-flight edits never land evenly.
What could the test have seen if the gain were real? "It was significant, so it was real." At 500 per arm on 3%, the chance of seeing it runs 5% to 24%.

Sources: Evan Miller (2010) for what stopping early costs, and Instantly's and Apollo's own help pages (checked August 2026) for what the two platforms do. Every percentage in the table is ours, computed from the inputs printed beside it.


Mechanism one

Looking at it every day is the whole problem

Opening the campaign on a Tuesday morning is the most normal thing a founder does.

The significance figure on your dashboard is only true if you fixed the size of the test before the send. Evan Miller named that assumption in 2010: the calculation "makes a critical assumption that you have probably violated without even realizing it: that the sample size was fixed in advance."

Four looking habits
1
One look

Set the stopping line before you send

One reading, at a point fixed before the send. That is the 5.0% you signed up for, and the row proves the simulation behaves.

4
Weekly

Drop the Monday morning review

Four weeks of it, stopping the moment the test reads significant, lands at 12.9%. The dashboard's 5% is already understated by more than double.

10
Every few days

Stop opening it every few days

Ten looks takes it to 19.4%, so one dead-even test in five hands you a winner. Miller's correction table agrees: at ten peeks, 1% is really 5%.

28
Daily, four weeks

Stop deciding on the morning refresh

One look a day for four weeks, and 25.6% of tests name a winner where the two arms are identical by construction. The habit, priced.

25.6%
Of tests with no real difference

report a significant winner, when the arms are identical and someone checks every morning for four weeks and stops at the first significant look.

Simulation, computed by us

Two arms of 1,000 contacts, both running at a 3% reply rate, with no real difference between them.

A low reply rate does not soften this, and neither does volume. Miller's worst case, at a 50% reply rate, lands at 26.1%. The work behind Optimizely's own statistics engine found that even at 10,000 contacts a side, peeking runs the false-winner rate to roughly five times what the dashboard prints.

Operator note
What I stopped doing

Refreshing the dashboard before coffee is not carelessness, it is caring, and I still do it. What I stopped doing is deciding on what I saw. We write the stop line before the send, and where a number is too thin to carry a claim, we leave the claim off.

KM
Kshitij Maheshwari
Co-founder, Real Good GTM

Mechanism two

Four variants is six comparisons

Five percent is what one pre-committed comparison costs, and nobody makes one comparison.

What you ran Comparisons Chance one looks significant when nothing is happening
Two variants, one metric, read once 1 5.0%
Three variants 3 14.3%
Four variants 6 26.5%
Three metrics read across four segments 12 46.0%
Two arithmetics, not one

That column is one minus 0.95 raised to the number of comparisons, computed by us, and four variants make six pairs rather than four. It holds only where comparisons are independent. Repeated looks at one running test are not, so ten looks is about 19%, never 40%.

So the move is two variants and one metric, both named before the send. If you also want to read the result by segment, that is a second test with its own list, not a second look at this one.


Mechanism three

The split you configured is not the split you got

Experimentation teams void a run over an unexpected split. At your volume you cannot detect one.

Rule, blind spot, fix
1
The rule

An off split usually voids the result

The experimentation teams at four software companies wrote this down together in 2019: a split that lands off target "in most cases completely invalidates experiment results". They throw the run away.

2
The blind spot

The check cannot see a test your size

Run their own check on 1,000 contacts and it only fires past about 53 against 47. A 52/48 split, plenty to bend a result at your volume, sails straight through it.

3
The fix

Build the split, do not test it

Shuffle the file before splitting it, split inside each mailbox and each send day rather than across them, and count both arms before the first send.

The split was rarely a coin flip to begin with: lists arrive in scrape order, one arm lands on a newer domain, the halves go out into different deliverability weather. HubSpot documents its own version of this: sends "may not be 50/50 (especially with smaller lists)".


Mechanism four

Two tools re-randomize at every step

If you tested two sequences rather than two emails, read your platform's A/B page first.

What you think you tested

"Sequence A beat sequence B by two points, so we are rolling out sequence A."

  • Assumes one lead sees one version end to end
  • Assumes step two belongs to the same arm as step one
  • Nothing in the dashboard says otherwise
What the documentation says

"A lead who receives Variant A in Step 1 could receive any variant (A, B, C, etc.) in Step 2. This is by design."

  • Instantly, on its own A/Z testing page, checked August 2026
  • Apollo documents the same behavior, updated 22 July 2026
  • So the unit you tested was the step, never the sequence

Neither hides it: both publish the behavior on their own help pages. What no dashboard does is size the test, and 100 or 200 delivered messages is a habit rather than a sample size, so compute the number before you send.

Instantly publishes a second fact worth knowing: it balances variant usage across a campaign's whole life, so add a variant mid-campaign and "only the new variants will be sent until their usage catches up".

Sources: Instantly Help Center, A/Z Testing (undated, checked August 2026); Apollo Knowledge Base (22 July 2026).

Operator note
Read the help page

The most useful three minutes in a testing week go on your platform's own A/B help page. Every founder can do it and almost nobody has. The rule it prints is rarely the rule you assumed the dashboard was following.

RB
Rahul Bageria
Co-founder, Real Good GTM

Want outbound run by two people who write the decision down first?

Book a Fit Check

Mechanism five

Your denominator moved while you watched

A rate has two moving parts and only one is replies. The question is not which denominator to pick, it is that yours did not hold still.

Don't

Quote the rate the dashboard hands you

Variant A 4.2% / Variant B 2.8% / Winner: A

  • No record of what sat on the bottom
  • Bounces left the arms different sizes
  • One thread counted twice on one arm only
Do

Write the whole fraction on the day you call it

A: 21 people replied of 500 enrolled on 3 Aug, 12 bounced, 2 opted out, each person counted once

  • Top and bottom fixed on one date
  • The same counting rule on both arms
  • A mid-flight edit shows up instead of hiding

An illustrative walkthrough of the mechanism, not a specific client result. We report real numbers only when they are real. The example numbers are invented.


Mechanism six

A test this small mostly cannot see anything

Before asking whether the winner is real, ask what the test could have found if it were.

Four parts, inputs printed
A worked case

7 replies against 12, at 100 sends each

A result shaped like that gets called a clear winner and rebuilt around. It is not one. At 100 sends a side, a gap that size turns up regularly when the two emails are identical.

The design

500 a side, at a 3% reply rate

Every figure below assumes two arms of 500 contacts and a 3.0% reply rate on the control. The arithmetic is ours, and it is the ordinary version of it.

What it sees

It misses almost any gain worth chasing

Against a real 5% improvement, the chance it spots the gain is 5.2%. At 20% it is 8.3%, at 50% it is 23.9%. Under a 15% improvement it never clears 10%.

What to trust

Read what the test could see, not what it saw

Numbers like these swing on a guess nobody can check in advance. At a real 10% improvement, 21.9% of the winners that clear the bar are the worse arm. At 5% it is 34.5%.

That chance of spotting a real gain is the number to lead with, because it stays low across every improvement anyone would plausibly claim. Two things go wrong underneath it, and both of them flatter you.

A test this thin can hand you a real winner dressed as the loser. And any winner that does clear the bar looks bigger than it is.

Five hundred a side is nowhere near enough to keep either of those off you. Separating a 3% reply rate from a 4.5% one takes about 2,500 contacts per variant. The volume table behind that price is on our guide to outbound market learning. That is the argument on that page. This is what follows from it.


The check

The winner that shrinks when you run it again

Everything above is diagnosis. This is the one thing here you can test in a week.

The prediction

Run the winner against the loser once more, on a fresh slice, re-randomized, and expect the gap to shrink. A margin that cleared the bar from that position is mostly the luck that got it there.

Re-randomize rather than reusing the same two halves. Reuse them and whatever the first test did to those contacts, who opened it, who was already mid-thread with you, rides into the second one and reads back as a result.

The same discipline applied to audiences rather than wording is our play on ICP slice experiments, which confirms a slice with a second batch.


The free diagnostic

Run one A/A test before you trust the next one

Spend one campaign finding out what your tool reports when there is nothing to report.

  1. 1

    Send the identical email to both halves

    Same copy, same subject line, same mailboxes, same send days. Label them A and B exactly as you would for a real test.

  2. 2

    Run it the way you run a real one

    Same volume, same duration, same number of morning check-ins. You are testing your procedure as much as your tool.

  3. 3

    Read the gap as your noise floor

    Whatever it reports came from your split, your mailboxes or your send days, because the copy was identical. Future winners have to beat it.

Ron Kohavi said it plainly back in 2012: "The A/A test has been our most useful tool in identifying issues in practical systems." One campaign is the whole price.


Failure modes

Five ways a result lies to you

The shapes those mechanisms take by the time they reach you, four of them dressed as good news.

The segment you found afterwards

It lost overall but won with ops leaders, so the test "worked". A segment chosen after reading the data is another comparison.

The metric that happened to move

You had reply rate, positive reply rate and meetings booked. Leading with the one that moved costs the same as running three tests.

"Observed power of 97%"

Worked out after the result is in, that number restates the result in different clothes. It cannot tell you whether the test was ever able to see what you are claiming.

A lift far too big to be true

Triple-digit lifts get published off a handful of replies on each side. The thinner the test, the more of the gap is the luck that produced it, and the less of it survives a re-run.

!
Caution

Auto-optimize will switch an arm off for you

Instantly documents that the feature "automatically analyzes variant performance and deactivates lower-performing versions, based on a defined winning metric", and open rate is one of the choices. A deactivated arm stops collecting replies, so the call cannot be revisited.

Do this instead
Leave it off until the test ends, or set the metric to replies, which is on the same list. An open is not a measurement.

What to do with all of that

A result is not a measurement until you can count the decisions that produced it.

Key takeaways
4 points
  • 1 Fix the sample size before the send, then read the result once.
  • 2 Two variants, one metric. Every extra comparison buys another false winner.
  • 3 Count both arms yourself. The split check is blind at your volume.
  • 4 Re-run the winner on a fresh split before you rebuild around it.

FAQ

Questions founders ask

How many emails do you need per variant for a cold email A/B test?
More than most tools suggest. At a 3% reply rate, separating 3% from 4.5% takes about 2,500 contacts per variant, and detecting a doubling to 6% takes about 750. At 100 per variant, the smallest difference you could reliably detect is roughly 3% against 14%.
Is it bad to stop a test early when one version is clearly winning?
Yes, and it is the most expensive habit on this page. The significance figure a dashboard prints assumes the sample size was fixed before the send. Checking daily for four weeks and stopping at the first significant reading turns a 5% false-positive rate into 25.6%, in a simulation where the arms are identical.
How many variants should I run at once?
Two. Four variants create six pairwise comparisons and a 26.5% chance at least one looks like a winner when nothing is happening: one minus 0.95 raised to the power six. At seed volume you also cannot split a list four ways and still read any arm.
Should I trust the winner my sending tool picked?
Read its documentation first, which takes about three minutes. Platforms differ on what declares a winner, how much of your list the test runs against, and whether a variant is carried through a sequence. Instantly and Apollo both document that assignment is random at every step.
What is an A/A test and why would I run one?
You send the identical email to both halves of a split and read what your tool reports. Any difference is your noise floor, because there is nothing else there. Ron Kohavi called it his most useful diagnostic back in 2012, and it costs you one campaign.
How do I tell a real winner from a lucky one after the fact?
Run the six checks at the top of this page, then re-run the test on a freshly randomized split. Real effects usually survive with the gap roughly intact, and lucky ones shrink toward zero. Re-randomize rather than reusing the same halves.
Kshitij Maheshwari, co-founder of Real Good GTM
About the author
Kshitij Maheshwari

Co-founder of Real Good GTM. He has been the first business hire and Chief of Staff at seed-stage B2B startups, building outbound pipeline before any playbook existed. He wrote this one because the argument he has most often with founders is about a winner that was never there, and the arithmetic settles it faster than opinion does.

Connect on LinkedIn

Keep going

Before you run the next one

The design side of the problem, the arithmetic behind the volumes, and the play that re-runs its own results.

Want outbound where the numbers survive a second look?

Book a fit check. We'll look at what your current numbers can and cannot support, set the one decision each campaign is allowed to make, and tell you straight if outbound is not the right motion for you yet.

Book a Fit Check

No hard sell. No fake numbers. Real good work speaks for itself.