Skip to content
GTM Guides

Cold email A/B testing: what to randomize

Most cold email tests are decided before the first send, by a choice nobody writes down: what is being randomized. How much volume a real test needs is settled and lives elsewhere on this site. This guide is about everything else you control.

By Kshitij Maheshwari, co-founder · Updated August 2026 · 22 min read

The short answer Five lines, then the detail
The first decision
Not what you are testing. What you are randomizing: a contact, a company, a mailbox or a segment. One word, chosen before the campaign is built.
Why it decides everything
Three contacts at one company are close to one observation, so every published sends-per-variant figure is a floor rather than a target.
What the web never taught you
Mailbox, sending domain, send hour, list order, warmup state and sequence step all differ between your two arms unless you stop them.
The artifact
A five-field card: the unit, the hold list, the metric, the decision rule, and what you do if there is no difference.
When the math says no
Eliminate rather than optimize, swing big, say directional out loud, and send one A/A test so you know what your own noise looks like.

Written by operators who run this motion for seed-stage teams, not by a tool that ships an A/B toggle.


What A/B testing means in outbound, and what it does not

The comparison is the easy half. The hard half is assignment: who ended up in which arm, and what else differed between the two groups without you deciding it.

Definition

A cold email A/B test runs two versions of one thing against comparably assigned groups from the same list, so a difference in replies belongs to the change rather than to who received it.

Also called split testing · full glossary

Three things separate this from a web experiment:

The population is finite

You are not sampling a stream of visitors. You are spending a list you paid to build, and it does not refill on its own.

Every unit burns on contact

A contact can be used once. Send them the losing version and they are not neutral ground for the next test.

The recipients remember

A person who read your first email reads the second one differently. Software has no memory of a treatment. Buyers do.

Everything a web experiment gets for free, an outbound test has to buy: units, time and a clean population. Which makes the design decisions matter more here, not less.


The volume question, settled in one paragraph

This part is already worked out, so it gets one paragraph and a link.

What a real test costs

At 500 sends a month, doubling a 2% positive reply rate takes about 1,141 contacts per variant and a 50% lift about 3,825. The assumptions are printed beside those figures in how to test an offer at seed volume, and this guide starts where that one stops.


One variable, changed all the way

Two halves of one rule, and most teams keep only the first. Follow that half alone and you get a perfectly isolated test of a change too small to see.

One variable, changed all the way. One so you can attribute it, all the way so you can see it.

The attribution half

Swap the slice, the opener and the channel together and no result can be assigned to any of them. Our ICP slice experiments play states that half plainly, and nothing here softens it.

The detection half

The one thing you moved has to move far enough to show up at your volume: test angles, not adjectives.

Both obey the one-variable rule exactly. "Quick question" against "Quick q" is cleanly isolated and undetectable at any volume a seed team will send. A benchmark offer against a teardown is also one variable, and big enough to land in the replies.


The unit

The first question is what you are randomizing

Four units are available in outbound, and each one changes what the test is able to conclude. The choice is made when you build the campaign, whether you make it on purpose or not.

The unit What you are asking What your sample really is What breaks
Contact Does this message work better on this person. The number of companies, once clustering is counted. Two people at one account get two different pitches from you. They may talk. Their gateway sees both.
Company Does this message work better on this account. The number of companies. You have to analyze per company too, or you overstate your own precision.
Mailbox or domain Does this mailbox perform better. The number of mailboxes, which is single digits. Sender and message are perfectly confounded. This is the most common accidental design in outbound.
Segment Does this message work better on this kind of company. The number of segments. Segment and message move together, so you learn nothing about either one.

The rule underneath the table: randomize where the version is independently applied, and read the result at that same level. A unit has to hold one version and stay independent of every other unit. Two contacts at one company fail the independence half.


Three contacts at one company are not three data points

Three people at one account are worth roughly one. Every sends-per-variant figure you have read is therefore a floor, and the way to clear it is more companies, not more people at each one.

What the count hides
1
The mechanism

One gateway, one budget, one conversation

People inside an account share a mail server, a buying situation and sometimes a thread about the email you just sent. Three replies from one company are three readings of the same company.

2
What it costs

The dashboard counts contacts, the test got companies

Sends per version is the number on the screen. Companies reached is the number the comparison runs on, and nobody has measured how far apart those two sit for cold email.

3
The consequence

Published sample sizes are floors, not targets

Every sends-per-variant figure printed anywhere, the table at the bottom of this page included, assumes one contact per company. Multi-thread three deep and the real requirement goes up from there.

4
The move

Buy accounts, not more seats at the same account

A hundred companies at one contact each is a sharper test than three hundred contacts spread across thirty companies. The next section turns that into a procedure you can run in a spreadsheet.


Randomize by company, analyze by company

The fix costs nothing, and it is the top row of the table above: fewer contacts per company inside a test, more companies in it.

  1. 1

    Assign whole companies, not people

    One coin flip per account, and every contact at that account inherits it. If your tool only splits contacts, do the assignment in the spreadsheet before the import.

  2. 2

    Keep one contact per company inside the test

    Adding companies buys you more than adding people at the companies you already have. The same total contacts, spread across more accounts, is the sharper test every time.

  3. 3

    Multi-thread the winner, not the test

    Contacting three people per account is right for booking meetings and wrong for learning. You do not have to choose forever. You do have to choose per campaign.


The confounds a web experiment never has

None of these exist when you split traffic on a web page, which is why the textbooks do not warn you about them and the sending tools do not either.

The confound How it gets in The design fix
Mailbox and domain Version A goes from one mailbox, version B from another. Reputation, warmup age and placement all differ. Every version sends from every mailbox in the pool, in the same proportions. Never a mailbox per version.
Send day and hour The tool works the list in order, so one version lands Tuesday morning and the other Friday afternoon. Randomize assignment inside each sending day, not across days.
List order Exports arrive sorted: alphabetically, by revenue, by whatever finished enriching last. Alternating rows assigns the sort. Shuffle the file before you split it. The next section is the whole procedure.
Warmup and volume state The campaign ramps. Early sends and late sends are not the same experiment. Hold daily volume flat while a test runs, or spread both versions evenly across the ramp.
Sequence step Step one and step four have different base rates, and on some tools the version does not follow a contact between them. Run the test on one step. Turn the other steps' versions off for the duration.
An outage or a block One version was in flight when a provider deferred your mail or a mailbox paused overnight. Nothing prevents this one. The split check further down is what catches it.
The rule that covers all six

Fisher wrote the general form in 1935, in The Design of Experiments: any later cause of difference that sits under your control must either be fixed before the versions are assigned, or randomized on its own account. Ninety-one years later, that one instruction is the whole of this section.


Shuffle the file

The cheapest correctness fix in this whole discipline, it takes one sort, and almost nobody does it.

Don't

Split the export the way it arrived

rows 1 to 250 = version A, rows 251 to 500 = version B

  • Every export is sorted by something
  • Alternating rows assigns that sort, not a coin
  • Your two arms differ before you send
Do

Shuffle, then split, then check

add a random column, sort by it, take the top half

  • Thirty seconds, no tooling, in the sheet
  • Shuffle the company list if you multi-thread
  • Keep the file so the split can be re-read

The artifact

The pre-send card

One page, five fields, written before the first send. If you cannot fill it in, you do not have a test yet. You have a send.

The five fields
The unit

Contact, company, mailbox or segment

One word. Everything below it depends on this answer, and it is the field no sending tool has ever asked you for.

The hold list

What stays still while this runs

Mailbox pool, sending window, daily volume, sequence length, list source, offer, ICP filter. Anything not on the list is either randomized on its own account or a confound you accepted.

The metric

One number, with its denominator named

Choose the metric and the bottom of the fraction before you can see either. Positive replies per contact is the honest target and the slowest to accumulate.

The decision rule

The difference you would act on

Plus the date you stop looking. Both are written before the send, because afterwards the rule quietly assembles itself around whatever arrived.

The three outcomes

A wins, B wins, no difference

Write what you will do in each case. The third row is the one everybody leaves blank, and at seed volume it is the most likely outcome by a distance.

Operator note
The field nobody fills

The third outcome is the one I have to argue for every time. A team that has not decided in advance what a tie means will keep whichever version they wrote first, and then believe they learned something.

KM
Kshitij Maheshwari
Co-founder, Real Good GTM

Why writing it down beats meaning to

Not because you would cheat. Because once the numbers arrive, a prediction and an explanation feel identical from the inside.

Why paper wins
1
The mechanism

Hindsight, not bad faith

Once you have seen the number, the reason for it turns up alongside it, fully formed, feeling exactly like something you knew on Monday. Nobody catches themselves doing this, which is the whole problem.

2
The fix

It is procedural, so willpower is not required

Write the question, the metric and the stop date while you still cannot see any of them. That is the pre-send card above, in a shared doc, and filling it in is the whole intervention.

3
The price

Ten minutes, and one uncomfortable decision

You have to say what would change your mind while you still have nothing invested in the answer. Do it before anyone opens the campaign builder, because afterwards it is a different decision.


Check that the split actually happened

Configuring a 50/50 split is not the same as delivering one, and outbound has more ways of breaking it than a web page does.

The five-minute check
1
What to count

Sends delivered, not contacts assigned

Count per version at the end of the test. If you configured an even split and you have 412 against 388, find out why before you read anything else on the dashboard.

2
Where it comes from

Bounces, filtering, a paused mailbox, a mid-flight edit

None of those land evenly across two versions. Neither does a version added late on a platform that balances usage across a campaign's whole lifetime.

3
The nuance

An uneven split is not a bug

Random assignment does produce uneven splits on small lists, and HubSpot's documentation says so. Check by eye that the two arms still match on seniority and company size. If one arm got the enterprise accounts, the test is void whatever the counts say.

Microsoft calls this a sample ratio mismatch, and its experimentation team reported in 2019 that roughly 6% of the experiments running an automated check hit one. They treat it as voiding a result, not weakening it. Counting your sends is your version of that check.


Pushback

One thing at a time is a concession, not a law

The rule is right for outbound. The reason usually given for it is not, and the difference decides how far you are willing to push the one thing you moved.

How the rule gets stated

"Change one variable at a time, because that is what statistical rigor requires."

  • Presented as a requirement of statistics
  • Implies two variables are unreadable at any scale
  • Silent on how far the one variable moves
  • Leaves a clean test of a change nobody sees
What it actually rests on

"If they have to be tested one at a time this is not because to do so is an ideal scientific procedure, but because to test them simultaneously would sometimes be too troublesome, or too costly."

  • R. A. Fisher, The Design of Experiments, 1935
  • Two variables are readable, given enough units
  • Outbound has none: each contact is spent once
  • So keep the rule, for attribution not rigor

The practical difference shows up when an arm breaks. On the web you rerun it next week. On a finite list those contacts are gone, so you keep the design simple to keep it debuggable.

Want a second read on the test you are about to run?

Book a Fit Check

What a holdout costs, and what it is actually for

Holding contacts back is cheap at seed volume and buys nothing, which is the opposite of the problem most teams expect.

At 500 sends a month and a 3% reply rate

Hold back 10% and you withhold 50 contacts a month, forgoing about 1.5 replies, or 150 contacts and roughly 4.5 replies a quarter. The cost is trivial. So is the return: 150 contacts cannot estimate anything to a useful precision.

Operator note
A reserve, not a control

We stopped calling our reserve a control, because at this volume it never behaved like one. What it actually is: the only clean list left for the next test, so we hold it back by company and keep whole accounts untouched.

RB
Rahul Bageria
Co-founder, Real Good GTM

One holdout does earn its keep, and it is a different animal: a segment you deliberately leave uncontacted for a quarter, so that a change in your reply rate can be told apart from a change in your market.


Carryover: the people in test two already got test one

Every guide tells you to run tests in sequence. None of them mentions that the second test runs on people who already met the first one.

Measured, then translated
1
How long

About three weeks, and once three months

Microsoft's experimentation team ran an experiment for 47 days, then kept watching the same groups of users after it stopped. The effect faded around the third week. Where a bug had given users a bad experience, those groups had not recovered three months on.

2
Why yours is worse

Their fix is not available to you

Their 2012 write-up blames an assignment system that never reshuffles people between experiments. You cannot reshuffle either: the people on your list read your last email and remember it.

3
What it costs

Fresh contacts, not ideas, set the cadence

A testing program burns list supply at the same rate as a sending program. Teams plan the tests and not the list, then run out of people for the third one.


The honest fallback

When you cannot power a test at all

This is where most readers actually stand: the calculator returned a number larger than a year of sending capacity. Six moves that still work, in order.

  1. 1

    Eliminate instead of optimize

    Zero positive replies in 300 well-targeted sends puts the 95% upper bound on that offer's true rate at 1%, by the rule of three. Killing is a claim your volume can support. Ranking two live options is not.

  2. 2

    Swing big or do not swing

    If the difference you are hunting is smaller than the difference your sample can see, the exercise is a lottery with a printout attached. Two sections down is the threshold for your volume.

  3. 3

    Run serially, knowing what serial does not buy

    For the same calendar time each arm gets the same sends either way: capacity times months, halved. Serial buys early exit, a simpler month, and the chance to design the second candidate after reading the first one's replies. It costs the time controls a split gives free.

  4. 4

    Read the replies as sentences

    A handful of replies is not a rate and is still plenty of evidence. That is rung two of our outbound market learning guide, and at your volume it is the read that pays.

  5. 5

    Say directional out loud, and mean it

    Directional means we changed the thing, the number moved, we cannot rule out chance, and we are acting anyway because acting costs less than waiting. That is a defensible decision. Calling it a winner is not.

  6. 6

    Send one A/A test

    The next section, and the only item here that tells you how large a difference your own tooling can produce out of nothing at all. It costs one campaign.

Operator note
Order of operations

Elimination is emotionally harder than optimization, and at this volume it is the only one that works. Optimizing feels like progress and produces a number. Eliminating feels like losing and produces a decision.

KM
Kshitij Maheshwari
Co-founder, Real Good GTM

Run one A/A test

Split the list as if you were testing, then send both halves the identical email. Whatever gap appears is your own noise, on your own list.

One campaign, four steps

4 checks

  • Split it as you would for a real test

    Same unit, same shuffle, same mailbox rotation. If the split is wrong here, it was going to be wrong on the test that mattered.

  • Send both halves the same email

    There is no second version to write. The whole cost is the discipline of changing nothing at all.

  • Count the sends, then read the gap

    Run the split check first. Then write down the difference between two identical emails, because that is what your dashboard reports from nothing at all.

  • Compare it to the last winner you called

    If the gap between two identical emails is as big as the win you declared last quarter, you have your answer and it cost you one campaign.

Microsoft's experimentation team calls the A/A test its most useful tool for finding problems in a live system, and runs them continuously. It is standard practice there and appears nowhere in cold email content.


How we would run it

The standing test calendar

An illustrative walkthrough of the mechanism, not a specific client result. We report real numbers only when they are real.

  1. 1

    Monday: write the card, before the builder opens

    Unit, hold list, metric, decision rule, and what happens on each of the three outcomes. Nobody touches the campaign until it exists.

  2. 2

    Weeks 1 to 12: one step, one sequence

    Volume stays flat, the mailbox pool stays fixed, the other steps' versions stay off. Nothing else ships while the test is live.

  3. 3

    The stop date: count first, read second

    It went in the calendar before the first send. On the day: sends per version, then the metric, then the card you wrote in week zero.

  4. 4

    Between tests: the next test needs fresh people

    Carryover means test two cannot run on the people who met test one. List supply, rather than the idea queue, sets the pace.

A test needing about 750 contacts per arm, at 500 sends a month, is a three-month test. Three or four a year is the honest number for a two-person team, and it forces the ideas to be big.


What works now

What the platforms actually do

Read from each vendor's own documentation and checked in August 2026. These are product behaviors, so they move, and the defaults matter more than the settings.

Platform How it splits Picks a winner for you? Winning metric options
Instantly Up to 26 versions per step, balanced across the campaign's whole lifetime, so a version added late is the only one sent until it catches up. Yes. Auto-optimize deactivates lower performers. Reply rate, click rate, open rate.
Smartlead Equal split, manual percentages, or an AI mode that moves sends toward the leader while the test is still running. Yes, in AI mode. Reply rate, positive reply rate, click rate, open rate.
Apollo Evenly per step, and a version cannot be pinned to a contact across steps. No. You deactivate underperformers yourself. Delivery and interested rates, spam-blocked and opt-out rates.
lemlist Splits across two sequence versions, and choosing a winner is permanent: leads already in the campaign keep their version, only new leads get the winner. No, and the choice cannot be undone. Compares open, reply and click.
HubSpot sequences Up to 6 versions per email, 4 live at once. The docs note that random assignment gives uneven splits on small lists. No. Sends, opens, clicks, replies, meetings booked.
The default that voids a multi-step test

Two of them document the same behavior independently: a contact who receives version A at step one can receive any version at step two. Instantly calls it by design; Apollo says a version cannot be assigned to a specific contact.

So a test spanning two steps leaves nobody holding a single treatment. Test one step, or switch the other steps' versions off.

Instantly and Smartlead will both pick a winner on open rate if you set it, and the others put opens beside replies for you to judge. After Apple's Mail Privacy Protection that is a setting to change, not a default to accept, and our subject line post carries that argument.


What 300 sends per variant can actually resolve

Turn the sample-size question around. Rather than asking what a test needs, ask what the sends you already have are able to see.

Contacts per version On a 3% reply rate, the smallest rate you can catch As a lift On a 2% reply rate As a lift
100 14.0% 366% 12.1% 504%
200 9.9% 228% 8.1% 306%
300 8.3% 175% 6.6% 232%
750 6.0% 100% 4.6% 129%
1,500 5.0% 67% 3.7% 85%
2,500 4.5% 50% 3.3% 63%

Computed by us and reproducible on any public sample size calculator at its usual settings: a real difference caught four times in five, a false winner one time in twenty. Read backwards, it answers what your sends can see. Run forwards, it returns about 750 per version to separate 3% from 6%, and about 2,500 to separate 3% from 4.5%.

Read the 300 row

At 300 contacts per version and a 3% reply rate, the smallest difference you have an 80% chance of catching is 3% against 8.3%. Anything narrower and you will miss it more often than you find it. Every row assumes one contact per company, so multi-threading pushes the requirement up again.

That is not an argument with anybody. It is a fact about your quarter.


What a test cannot fix

Design discipline makes a comparison readable. It does nothing for either side of the comparison.

The offer is upstream

If the offer is wrong, a perfectly designed test tells you precisely which of two wrong messages is less wrong. That is a real answer to a question you should not be asking yet.

So is the list

A clean split across the wrong accounts still produces a clean number. Randomization protects attribution. It has never once protected relevance.

The through-line

Run a test when you have a real question and a difference big enough to answer it. Everything on this page is about making the answer trustworthy. None of it is about making the answer good.


What to keep from all of this

If a test you are about to run breaks one of these, fix that before you fix the copy.

Key takeaways
5 points
  • 1 Name the unit before the copy: contact, company, mailbox or segment.
  • 2 Contacts at one company are close to one observation, not three.
  • 3 Hold the mailbox, the hour, the step and the list order still.
  • 4 Write the decision rule and the tie outcome before the first send.
  • 5 Count sends per version before you read a single rate.

One variable, changed all the way, on a unit you chose on purpose.


FAQ

Questions founders ask

What is cold email A/B testing?
Running two versions of one thing against comparably assigned groups from the same list, so a difference in replies belongs to the change rather than to who received it. The comparison is the easy half. The hard half is assignment: what you randomized, and what else differed between the two groups without you choosing it.
Should I split by contact or by company?
By company, if you contact more than one person per account. Contacts at the same employer share a mail gateway, a buying situation and sometimes a conversation, so they are not independent observations. Randomize whole companies to one version, analyze at the company level, and keep one contact per company inside the test itself.
How many emails do I need per variant?
More than most defaults suggest, and the honest way to see it is to turn the question around. At 100 contacts per version and a 3% reply rate you have three replies to compare against three replies, and the smallest gap you could reliably catch is 3% against 14%. At 300 per version it is 3% against 8.3%.
Why does my winner disappear when I run the test again?
Usually because the first result was noise, and sometimes because the second test ran on the people who received the first one. Microsoft's experimentation team measured that hangover lasting about three weeks after an experiment ended, and in one case more than three months. Fresh contacts per test is the only reliable fix in outbound.
What is an A/A test and why would I send the same email twice?
You split the list as if you were testing, then send both halves the identical email. Any gap you see is noise, produced by your own tools on your own list. It is a standard trustworthiness check at companies that run experiments for a living, it costs one campaign, and it is the fastest cure for acting on a two-point difference.
Should I keep a holdout group?
Not as a control. Withholding 10% of a 500-a-month program costs about 150 contacts a quarter and produces an estimate too imprecise to settle anything. The reserve is worth keeping for a different reason: contacts you have not emailed are the only clean population left for the next test, so hold them back by company.
My tool picked a winner automatically. Can I trust it?
Read what it optimized. Instantly and Smartlead both accept open rate as the winning metric, and Smartlead's AI mode moves sends toward the leader while the test is still running. That second one is peeking, which our post on why your A/B test is lying takes apart. Turn auto-optimize off for a test you intend to read, and check the default first.
Kshitij Maheshwari, co-founder of Real Good GTM
About the author
Kshitij Maheshwari

Co-founder of Real Good GTM. He has been the first business hire and Chief of Staff at seed-stage B2B startups, building outbound pipeline before any playbook existed. He wrote this guide because the tests he has watched fail at seed stage almost never failed on the arithmetic. They failed on a decision nobody wrote down.

Connect on LinkedIn

Keep going

From the design to the read

You have the design half. These three cover the rest: the forensics on a finished result, the arithmetic a real test costs, and what a rate can prove.

Want this run for you, tests and all?

Book a fit check. We'll look at your ICP, what you are actually trying to learn, and whether your sending volume can answer it this quarter. If it cannot, we'll tell you what will.

Book a Fit Check

No hard sell. No fake numbers. Real good work speaks for itself.