Cold email A/B testing: what to randomize
Most cold email tests are decided before the first send, by a choice nobody writes down: what is being randomized. How much volume a real test needs is settled and lives elsewhere on this site. This guide is about everything else you control.
By Kshitij Maheshwari, co-founder · Updated August 2026 · 22 min read
Written by operators who run this motion for seed-stage teams, not by a tool that ships an A/B toggle.
What A/B testing means in outbound, and what it does not
The comparison is the easy half. The hard half is assignment: who ended up in which arm, and what else differed between the two groups without you deciding it.
A cold email A/B test runs two versions of one thing against comparably assigned groups from the same list, so a difference in replies belongs to the change rather than to who received it.
Three things separate this from a web experiment:
You are not sampling a stream of visitors. You are spending a list you paid to build, and it does not refill on its own.
A contact can be used once. Send them the losing version and they are not neutral ground for the next test.
A person who read your first email reads the second one differently. Software has no memory of a treatment. Buyers do.
Everything a web experiment gets for free, an outbound test has to buy: units, time and a clean population. Which makes the design decisions matter more here, not less.
The volume question, settled in one paragraph
This part is already worked out, so it gets one paragraph and a link.
At 500 sends a month, doubling a 2% positive reply rate takes about 1,141 contacts per variant and a 50% lift about 3,825. The assumptions are printed beside those figures in how to test an offer at seed volume, and this guide starts where that one stops.
One variable, changed all the way
Two halves of one rule, and most teams keep only the first. Follow that half alone and you get a perfectly isolated test of a change too small to see.
One variable, changed all the way. One so you can attribute it, all the way so you can see it.
Swap the slice, the opener and the channel together and no result can be assigned to any of them. Our ICP slice experiments play states that half plainly, and nothing here softens it.
The one thing you moved has to move far enough to show up at your volume: test angles, not adjectives.
Both obey the one-variable rule exactly. "Quick question" against "Quick q" is cleanly isolated and undetectable at any volume a seed team will send. A benchmark offer against a teardown is also one variable, and big enough to land in the replies.
The first question is what you are randomizing
Four units are available in outbound, and each one changes what the test is able to conclude. The choice is made when you build the campaign, whether you make it on purpose or not.
| The unit | What you are asking | What your sample really is | What breaks |
|---|---|---|---|
| Contact | Does this message work better on this person. | The number of companies, once clustering is counted. | Two people at one account get two different pitches from you. They may talk. Their gateway sees both. |
| Company | Does this message work better on this account. | The number of companies. | You have to analyze per company too, or you overstate your own precision. |
| Mailbox or domain | Does this mailbox perform better. | The number of mailboxes, which is single digits. | Sender and message are perfectly confounded. This is the most common accidental design in outbound. |
| Segment | Does this message work better on this kind of company. | The number of segments. | Segment and message move together, so you learn nothing about either one. |
The rule underneath the table: randomize where the version is independently applied, and read the result at that same level. A unit has to hold one version and stay independent of every other unit. Two contacts at one company fail the independence half.
Three contacts at one company are not three data points
Three people at one account are worth roughly one. Every sends-per-variant figure you have read is therefore a floor, and the way to clear it is more companies, not more people at each one.
One gateway, one budget, one conversation
People inside an account share a mail server, a buying situation and sometimes a thread about the email you just sent. Three replies from one company are three readings of the same company.
The dashboard counts contacts, the test got companies
Sends per version is the number on the screen. Companies reached is the number the comparison runs on, and nobody has measured how far apart those two sit for cold email.
Published sample sizes are floors, not targets
Every sends-per-variant figure printed anywhere, the table at the bottom of this page included, assumes one contact per company. Multi-thread three deep and the real requirement goes up from there.
Buy accounts, not more seats at the same account
A hundred companies at one contact each is a sharper test than three hundred contacts spread across thirty companies. The next section turns that into a procedure you can run in a spreadsheet.
Randomize by company, analyze by company
The fix costs nothing, and it is the top row of the table above: fewer contacts per company inside a test, more companies in it.
-
1
Assign whole companies, not people
One coin flip per account, and every contact at that account inherits it. If your tool only splits contacts, do the assignment in the spreadsheet before the import.
-
2
Keep one contact per company inside the test
Adding companies buys you more than adding people at the companies you already have. The same total contacts, spread across more accounts, is the sharper test every time.
-
3
Multi-thread the winner, not the test
Contacting three people per account is right for booking meetings and wrong for learning. You do not have to choose forever. You do have to choose per campaign.
The confounds a web experiment never has
None of these exist when you split traffic on a web page, which is why the textbooks do not warn you about them and the sending tools do not either.
| The confound | How it gets in | The design fix |
|---|---|---|
| Mailbox and domain | Version A goes from one mailbox, version B from another. Reputation, warmup age and placement all differ. | Every version sends from every mailbox in the pool, in the same proportions. Never a mailbox per version. |
| Send day and hour | The tool works the list in order, so one version lands Tuesday morning and the other Friday afternoon. | Randomize assignment inside each sending day, not across days. |
| List order | Exports arrive sorted: alphabetically, by revenue, by whatever finished enriching last. Alternating rows assigns the sort. | Shuffle the file before you split it. The next section is the whole procedure. |
| Warmup and volume state | The campaign ramps. Early sends and late sends are not the same experiment. | Hold daily volume flat while a test runs, or spread both versions evenly across the ramp. |
| Sequence step | Step one and step four have different base rates, and on some tools the version does not follow a contact between them. | Run the test on one step. Turn the other steps' versions off for the duration. |
| An outage or a block | One version was in flight when a provider deferred your mail or a mailbox paused overnight. | Nothing prevents this one. The split check further down is what catches it. |
Fisher wrote the general form in 1935, in The Design of Experiments: any later cause of difference that sits under your control must either be fixed before the versions are assigned, or randomized on its own account. Ninety-one years later, that one instruction is the whole of this section.
Shuffle the file
The cheapest correctness fix in this whole discipline, it takes one sort, and almost nobody does it.
Split the export the way it arrived
rows 1 to 250 = version A, rows 251 to 500 = version B
- ✕Every export is sorted by something
- ✕Alternating rows assigns that sort, not a coin
- ✕Your two arms differ before you send
Shuffle, then split, then check
add a random column, sort by it, take the top half
- ✓Thirty seconds, no tooling, in the sheet
- ✓Shuffle the company list if you multi-thread
- ✓Keep the file so the split can be re-read
The pre-send card
One page, five fields, written before the first send. If you cannot fill it in, you do not have a test yet. You have a send.
Contact, company, mailbox or segment
One word. Everything below it depends on this answer, and it is the field no sending tool has ever asked you for.
What stays still while this runs
Mailbox pool, sending window, daily volume, sequence length, list source, offer, ICP filter. Anything not on the list is either randomized on its own account or a confound you accepted.
One number, with its denominator named
Choose the metric and the bottom of the fraction before you can see either. Positive replies per contact is the honest target and the slowest to accumulate.
The difference you would act on
Plus the date you stop looking. Both are written before the send, because afterwards the rule quietly assembles itself around whatever arrived.
A wins, B wins, no difference
Write what you will do in each case. The third row is the one everybody leaves blank, and at seed volume it is the most likely outcome by a distance.
The third outcome is the one I have to argue for every time. A team that has not decided in advance what a tie means will keep whichever version they wrote first, and then believe they learned something.
Why writing it down beats meaning to
Not because you would cheat. Because once the numbers arrive, a prediction and an explanation feel identical from the inside.
Hindsight, not bad faith
Once you have seen the number, the reason for it turns up alongside it, fully formed, feeling exactly like something you knew on Monday. Nobody catches themselves doing this, which is the whole problem.
It is procedural, so willpower is not required
Write the question, the metric and the stop date while you still cannot see any of them. That is the pre-send card above, in a shared doc, and filling it in is the whole intervention.
Ten minutes, and one uncomfortable decision
You have to say what would change your mind while you still have nothing invested in the answer. Do it before anyone opens the campaign builder, because afterwards it is a different decision.
Check that the split actually happened
Configuring a 50/50 split is not the same as delivering one, and outbound has more ways of breaking it than a web page does.
Sends delivered, not contacts assigned
Count per version at the end of the test. If you configured an even split and you have 412 against 388, find out why before you read anything else on the dashboard.
Bounces, filtering, a paused mailbox, a mid-flight edit
None of those land evenly across two versions. Neither does a version added late on a platform that balances usage across a campaign's whole lifetime.
An uneven split is not a bug
Random assignment does produce uneven splits on small lists, and HubSpot's documentation says so. Check by eye that the two arms still match on seniority and company size. If one arm got the enterprise accounts, the test is void whatever the counts say.
Microsoft calls this a sample ratio mismatch, and its experimentation team reported in 2019 that roughly 6% of the experiments running an automated check hit one. They treat it as voiding a result, not weakening it. Counting your sends is your version of that check.
One thing at a time is a concession, not a law
The rule is right for outbound. The reason usually given for it is not, and the difference decides how far you are willing to push the one thing you moved.
"Change one variable at a time, because that is what statistical rigor requires."
- ✕Presented as a requirement of statistics
- ✕Implies two variables are unreadable at any scale
- ✕Silent on how far the one variable moves
- ✕Leaves a clean test of a change nobody sees
"If they have to be tested one at a time this is not because to do so is an ideal scientific procedure, but because to test them simultaneously would sometimes be too troublesome, or too costly."
- ✓R. A. Fisher, The Design of Experiments, 1935
- ✓Two variables are readable, given enough units
- ✓Outbound has none: each contact is spent once
- ✓So keep the rule, for attribution not rigor
The practical difference shows up when an arm breaks. On the web you rerun it next week. On a finite list those contacts are gone, so you keep the design simple to keep it debuggable.
Want a second read on the test you are about to run?
Book a Fit CheckWhat a holdout costs, and what it is actually for
Holding contacts back is cheap at seed volume and buys nothing, which is the opposite of the problem most teams expect.
Hold back 10% and you withhold 50 contacts a month, forgoing about 1.5 replies, or 150 contacts and roughly 4.5 replies a quarter. The cost is trivial. So is the return: 150 contacts cannot estimate anything to a useful precision.
We stopped calling our reserve a control, because at this volume it never behaved like one. What it actually is: the only clean list left for the next test, so we hold it back by company and keep whole accounts untouched.
One holdout does earn its keep, and it is a different animal: a segment you deliberately leave uncontacted for a quarter, so that a change in your reply rate can be told apart from a change in your market.
Carryover: the people in test two already got test one
Every guide tells you to run tests in sequence. None of them mentions that the second test runs on people who already met the first one.
About three weeks, and once three months
Microsoft's experimentation team ran an experiment for 47 days, then kept watching the same groups of users after it stopped. The effect faded around the third week. Where a bug had given users a bad experience, those groups had not recovered three months on.
Their fix is not available to you
Their 2012 write-up blames an assignment system that never reshuffles people between experiments. You cannot reshuffle either: the people on your list read your last email and remember it.
Fresh contacts, not ideas, set the cadence
A testing program burns list supply at the same rate as a sending program. Teams plan the tests and not the list, then run out of people for the third one.
When you cannot power a test at all
This is where most readers actually stand: the calculator returned a number larger than a year of sending capacity. Six moves that still work, in order.
-
1
Eliminate instead of optimize
Zero positive replies in 300 well-targeted sends puts the 95% upper bound on that offer's true rate at 1%, by the rule of three. Killing is a claim your volume can support. Ranking two live options is not.
-
2
Swing big or do not swing
If the difference you are hunting is smaller than the difference your sample can see, the exercise is a lottery with a printout attached. Two sections down is the threshold for your volume.
-
3
Run serially, knowing what serial does not buy
For the same calendar time each arm gets the same sends either way: capacity times months, halved. Serial buys early exit, a simpler month, and the chance to design the second candidate after reading the first one's replies. It costs the time controls a split gives free.
-
4
Read the replies as sentences
A handful of replies is not a rate and is still plenty of evidence. That is rung two of our outbound market learning guide, and at your volume it is the read that pays.
-
5
Say directional out loud, and mean it
Directional means we changed the thing, the number moved, we cannot rule out chance, and we are acting anyway because acting costs less than waiting. That is a defensible decision. Calling it a winner is not.
-
6
Send one A/A test
The next section, and the only item here that tells you how large a difference your own tooling can produce out of nothing at all. It costs one campaign.
Elimination is emotionally harder than optimization, and at this volume it is the only one that works. Optimizing feels like progress and produces a number. Eliminating feels like losing and produces a decision.
Run one A/A test
Split the list as if you were testing, then send both halves the identical email. Whatever gap appears is your own noise, on your own list.
One campaign, four steps
4 checks
-
Split it as you would for a real test
Same unit, same shuffle, same mailbox rotation. If the split is wrong here, it was going to be wrong on the test that mattered.
-
Send both halves the same email
There is no second version to write. The whole cost is the discipline of changing nothing at all.
-
Count the sends, then read the gap
Run the split check first. Then write down the difference between two identical emails, because that is what your dashboard reports from nothing at all.
-
Compare it to the last winner you called
If the gap between two identical emails is as big as the win you declared last quarter, you have your answer and it cost you one campaign.
Microsoft's experimentation team calls the A/A test its most useful tool for finding problems in a live system, and runs them continuously. It is standard practice there and appears nowhere in cold email content.
The standing test calendar
An illustrative walkthrough of the mechanism, not a specific client result. We report real numbers only when they are real.
-
1
Monday: write the card, before the builder opens
Unit, hold list, metric, decision rule, and what happens on each of the three outcomes. Nobody touches the campaign until it exists.
-
2
Weeks 1 to 12: one step, one sequence
Volume stays flat, the mailbox pool stays fixed, the other steps' versions stay off. Nothing else ships while the test is live.
-
3
The stop date: count first, read second
It went in the calendar before the first send. On the day: sends per version, then the metric, then the card you wrote in week zero.
-
4
Between tests: the next test needs fresh people
Carryover means test two cannot run on the people who met test one. List supply, rather than the idea queue, sets the pace.
A test needing about 750 contacts per arm, at 500 sends a month, is a three-month test. Three or four a year is the honest number for a two-person team, and it forces the ideas to be big.
What the platforms actually do
Read from each vendor's own documentation and checked in August 2026. These are product behaviors, so they move, and the defaults matter more than the settings.
| Platform | How it splits | Picks a winner for you? | Winning metric options |
|---|---|---|---|
| Instantly | Up to 26 versions per step, balanced across the campaign's whole lifetime, so a version added late is the only one sent until it catches up. | Yes. Auto-optimize deactivates lower performers. | Reply rate, click rate, open rate. |
| Smartlead | Equal split, manual percentages, or an AI mode that moves sends toward the leader while the test is still running. | Yes, in AI mode. | Reply rate, positive reply rate, click rate, open rate. |
| Apollo | Evenly per step, and a version cannot be pinned to a contact across steps. | No. You deactivate underperformers yourself. | Delivery and interested rates, spam-blocked and opt-out rates. |
| lemlist | Splits across two sequence versions, and choosing a winner is permanent: leads already in the campaign keep their version, only new leads get the winner. | No, and the choice cannot be undone. | Compares open, reply and click. |
| HubSpot sequences | Up to 6 versions per email, 4 live at once. The docs note that random assignment gives uneven splits on small lists. | No. | Sends, opens, clicks, replies, meetings booked. |
Two of them document the same behavior independently: a contact who receives version A at step one can receive any version at step two. Instantly calls it by design; Apollo says a version cannot be assigned to a specific contact.
So a test spanning two steps leaves nobody holding a single treatment. Test one step, or switch the other steps' versions off.
Instantly and Smartlead will both pick a winner on open rate if you set it, and the others put opens beside replies for you to judge. After Apple's Mail Privacy Protection that is a setting to change, not a default to accept, and our subject line post carries that argument.
What 300 sends per variant can actually resolve
Turn the sample-size question around. Rather than asking what a test needs, ask what the sends you already have are able to see.
| Contacts per version | On a 3% reply rate, the smallest rate you can catch | As a lift | On a 2% reply rate | As a lift |
|---|---|---|---|---|
| 100 | 14.0% | 366% | 12.1% | 504% |
| 200 | 9.9% | 228% | 8.1% | 306% |
| 300 | 8.3% | 175% | 6.6% | 232% |
| 750 | 6.0% | 100% | 4.6% | 129% |
| 1,500 | 5.0% | 67% | 3.7% | 85% |
| 2,500 | 4.5% | 50% | 3.3% | 63% |
Computed by us and reproducible on any public sample size calculator at its usual settings: a real difference caught four times in five, a false winner one time in twenty. Read backwards, it answers what your sends can see. Run forwards, it returns about 750 per version to separate 3% from 6%, and about 2,500 to separate 3% from 4.5%.
At 300 contacts per version and a 3% reply rate, the smallest difference you have an 80% chance of catching is 3% against 8.3%. Anything narrower and you will miss it more often than you find it. Every row assumes one contact per company, so multi-threading pushes the requirement up again.
That is not an argument with anybody. It is a fact about your quarter.
What a test cannot fix
Design discipline makes a comparison readable. It does nothing for either side of the comparison.
If the offer is wrong, a perfectly designed test tells you precisely which of two wrong messages is less wrong. That is a real answer to a question you should not be asking yet.
A clean split across the wrong accounts still produces a clean number. Randomization protects attribution. It has never once protected relevance.
Run a test when you have a real question and a difference big enough to answer it. Everything on this page is about making the answer trustworthy. None of it is about making the answer good.
What to keep from all of this
If a test you are about to run breaks one of these, fix that before you fix the copy.
- 1 Name the unit before the copy: contact, company, mailbox or segment.
- 2 Contacts at one company are close to one observation, not three.
- 3 Hold the mailbox, the hour, the step and the list order still.
- 4 Write the decision rule and the tie outcome before the first send.
- 5 Count sends per version before you read a single rate.
One variable, changed all the way, on a unit you chose on purpose.
Questions founders ask
What is cold email A/B testing?
Should I split by contact or by company?
How many emails do I need per variant?
Why does my winner disappear when I run the test again?
What is an A/A test and why would I send the same email twice?
Should I keep a holdout group?
My tool picked a winner automatically. Can I trust it?
Co-founder of Real Good GTM. He has been the first business hire and Chief of Staff at seed-stage B2B startups, building outbound pipeline before any playbook existed. He wrote this guide because the tests he has watched fail at seed stage almost never failed on the arithmetic. They failed on a decision nobody wrote down.
Connect on LinkedInFrom the design to the read
You have the design half. These three cover the rest: the forensics on a finished result, the arithmetic a real test costs, and what a rate can prove.
Why your A/B test is lying
The forensics on a result you already believe: peeking, extra variants, a moving denominator, and the checks that catch each one.
Read the postThe outbound offer
What a real test costs at seed volume, the one bound you can honestly claim, and why whole offers beat two subject lines.
Read the guideTurning outbound into GTM learning
What a reply rate can and cannot prove at your volume, the evidence ladder, and the memo to write the findings down in.
Read the guideWant this run for you, tests and all?
Book a fit check. We'll look at your ICP, what you are actually trying to learn, and whether your sending volume can answer it this quarter. If it cannot, we'll tell you what will.
Book a Fit CheckNo hard sell. No fake numbers. Real good work speaks for itself.