How to test cold email offers
You test a cold email offer by running one candidate at a time into a frozen segment, against a rule you wrote before the first send. The part nobody writes down sets the price: how far apart the two candidates are is a decision you make for free, and it enters the cost of the test squared.
By Rahul Bageria, co-founder · Updated August 2026 · 13 min read
One offer test, seven steps
The whole protocol is one table. Note where the cost sits: the two steps that decide whether a test exists cost nothing.
| Step | What it produces | What it costs |
|---|---|---|
| 1. Pick two candidates | Two whole offers you would ship if either won. | An afternoon. No sends. |
| 2. Run the distance check | A yes or no on whether a comparison exists. | Ten minutes. No sends. |
| 3. Freeze the setup | One segment, one list build, one sequence, one window. | One list build, order randomized once. |
| 4. Write the kill rule | A number, a date, an action, and a name. | Ten minutes, before the first send. |
| 5. Send candidate one | Replies, sorted, and a count running against the rule. | A few hundred sends to kill it. Thousands to compare. |
| 6. Close on the date | Promote it, keep the incumbent, or retire it. | Nothing, if the date was set in advance. |
| 7. Write the log entry | Seven lines that make the next test cheaper. | Fifteen minutes on the day it closes. |
The send counts in step 5 come from our cold email offer guide. The next section quotes them and links to the working.
What the offer guide already settled
Two numbers govern the price of an offer test. Both are published, so this page quotes them once and moves on.
Detecting a doubling of positive reply rate takes about 1,141 sends per candidate. Killing one is the affordable verdict: after 300 well-targeted sends with no positive reply, the most that offer could plausibly be doing is a 1% reply rate, and you can retire it.
Both, with their method and sources, sit on our cold email offer guide. This page starts where that stops.
Distance is the only lever that costs nothing
Volume, copy and timing all cost you money or weeks. How far apart your two candidates sit costs an afternoon, and it sets the price of everything after it.
Halve the difference you are chasing and the sends roughly quadruple. Nothing else you control moves the bill that fast.
The 1,141 sends above buy you one question: does this offer double the reply rate? At 500 sends a month, that is about five months to run one test.
Ask a finer question, whether it lifts you from 2% to 2.5%, and the same guide prices it at about 13,809 sends per candidate. That is longer than most runways.
Same offer, same list, same month. You made the target four times finer, and the bill went up more than tenfold.
So the first question is not what to test. It is whether these two candidates sit far enough apart to be worth four weeks of a segment you cannot refill, and that check takes ten minutes and no sends.
How to tell whether two offers are far apart
Measure it on the four parts an offer already has, plus the time horizon. Near is different wording. Far is a different thing arriving.
| Axis | The question | Near: one offer, two drafts | Far: a real comparison |
|---|---|---|---|
| Claim | Which problem you name | The same problem, said more sharply. | A different problem, owned by the same person. |
| Substance | What they actually receive | The same document with a better title. | A document, against work done on their own thing. |
| Alternatives | What it beats | Both beat doing nothing. | One beats doing nothing, one beats the tool they bought. |
| The ask | What accepting costs them | "Fifteen minutes" against "a quick chat". | One word back, against access to their data. |
| Time horizon | When it pays off | This quarter, both of them. | Something by Friday, against a plan for next year. |
A sixth axis is a disqualifier, not a distance. If the candidates need different people to be true, you changed the audience, and that is a slice experiment, which holds the offer still and moves the audience.
This is the detection half of a rule our cold email A/B testing guide states in full: one variable, changed all the way.
Two offers are far enough apart when the same objection cannot kill both. If one sentence kills both, you have one candidate and two drafts.
Run it out loud, against candidate A then candidate B. The rule has an honest cost: push two far apart and one is often an offer you would never run, so both have to be offers you would ship.
Same container, same offer
The offer guide lists six shapes and calls the shape the container. That gives you a two-second version of the distance check.
- •Two benchmarks with different titles
- •Teardowns of two different pages
- •The same document, renamed for a different reader
- •One objection kills the pair, so the test is over already
- •A benchmark against a teardown
- •A short list against a migration angle
- •An introduction against a slice pilot
- •They fail differently, so no single sentence kills both
This is why the offer guide says to run a benchmark against a teardown against a migration.
What you freeze, and what you cannot
"Hold everything else constant" is an instruction everywhere and a list nowhere. Here is the list, and the longer one that moves anyway.
Freeze it in your head
Same segment, roughly. We'll keep the sequence the same and see how it goes.
- ✕"Roughly" is where the test dies
- ✕Nothing written down is nothing frozen
- ✕"See how it goes" has no end date
Freeze it in writing, before the build
Segment: US seed SaaS, 20 to 80 staff. One build, order randomized 12 Aug. Sequence: 3 steps, days 0, 3, 8. 4 mailboxes, 25 a day. Window: 4 whole weeks from 17 Aug. Positive = they ask for the thing.
- ✓Six lines, written before any build
- ✓Whole weeks, so send days land evenly
- ✓"Positive" defined before the first reply
Five things move regardless: list quality drift, mailbox reputation, the calendar, competitors sending the same angle into the same inboxes, and you.
Printing the second list is what makes the first credible. The costliest drift goes unnoticed: send to the best-fitting accounts first and candidate A ran against a better list. Randomize the whole list once, at build time, then draw in order.
No edits inside a test window. At day ten with four replies in, the instinct is to fix the subject line. Either it is too small to matter and you traded a clean record for nothing, or it is large and you are now on a third candidate.
Why one offer at a time
Start with the asymmetry, because it is the actual reason. A dead offer announces itself quickly. Two living offers separate slowly, and no arrangement of the volume changes that.
"Run both offers at once, fifty-fifty, and let the dashboard call the winner."
- ✕Commits both arms before the first reply lands
- ✕Needs both arms to reach thousands of sends
- ✕Below the volume line, resolves nothing at all
"Run candidate one to a date. Read the replies. Then design candidate two."
- ✓Early exit: kill it partway and reallocate the rest
- ✓Sequencing: candidate two is designed, not guessed
- ✓One list, one thing moving, no split to verify
None of those three is statistical, and here is the part most operators have backwards.
Running one at a time does not buy a bigger sample. Same capacity, same segment, same weeks: a split gives each arm half the volume, and back to back gives each candidate half. Identical sample either way.
The two candidates ran in different weeks, against different list freshness, a different deliverability state and a different market. A split controls all of that for free, so you can never say offer B beat offer A, only that each cleared the bar or did not.
The volume line, with its inputs: two arms at the doubling figure above, 1,141 sends each, is 2,282 sends. That is roughly 1,650 a month into one segment over six weeks, or 1,240 over eight. Below it, a split resolves nothing.
Want a second read on the two offers you were about to spend a month on?
Book a Fit CheckThe kill rule, written before the send
Five fields, one sentence, written before the first send, so the decision is not made later by whoever is most invested.
-
1
The metric, and its denominator
Positive replies per contact, counted once per person, not per thread. Write the definition before the first reply arrives, or you are choosing a rule that fits the replies you already have.
-
2
The number
The count that ends it. Zero positives is the clean one, because it needs no comparison. Any other threshold you should be able to defend before seeing the data.
-
3
The date
In whole weeks, set in advance. Kohavi's team recommend multiples of a week so day-of-week effects can be read, and a fixed date stops you stopping on a good number.
-
4
The action
What happens the morning after. Promote it, keep the incumbent, or retire the candidate. Naming it in advance stops a finished test from quietly continuing.
-
5
Who executes it
A name, and not whoever wrote the offer. On a two-person team the other founder runs the rule on the date, without a meeting about it.
The person who wrote the offer is the worst-placed person to call it. Not because anyone is dishonest, but because you can always find a reason to give your own idea one more week.
What closes a test that neither wins nor dies
Most offer tests end in the middle. Not zero positives, so nothing fired. Nowhere near a doubling, so nothing was proven either.
The scarce thing is the slot, not significance. At 500 sends a month a quarter is 1,500 sends and killing one candidate costs 300, so half the volume on the incumbent leaves room for two or three.
Two moves are off the table: extending the window because the number is close, and calling a winner on a difference smaller than you designed for. The first has a mechanism, in why your A/B test is lying.
- 1 The date closes the test. The number decides what happens next.
- 2 Stretching a window until it looks good is not a result.
- 3 Two or three candidates a quarter is the real budget.
- 4 Ask whether it deserves the slot, not whether it won.
The log entry
Seven lines, written the day the test closes. Without them, test three starts from a guess, exactly like test one.
The window as it ran, not as planned
Start, end, and whether it ran in whole weeks. A test that stopped early for any reason but the kill rule is a different record.
As built, with the build date
Filters, size, source, and the day the list was built. Two tests into one named segment are not comparable if a list was six weeks older.
One sentence, plus which axis moved
The candidate in a sentence, and which axis moved against the incumbent: claim, substance, alternatives, ask or time horizon. The distance, recorded.
With the denominator named
Contacts, delivered, replies, positives, and what counted as positive. The denominator is part of the number, and you will not remember which one next quarter.
What it said, and whether it fired
The rule word for word, and the outcome. One that did not fire is as useful as one that did. One quietly rewritten is the most useful line here.
Who made it, on what date
Promote, keep or retire, with a name and a date on it. The decision is the artifact; the rate is a supporting detail, and a noisy one.
What came back most
The one sentence that showed up most often. It tells you which axis to move next, and it is the cheapest market research you will get.
One entry is a record. A quarter of them is a read on your market, the subject of our guide to turning outbound into GTM learning.
One offer test on a calendar
An illustrative walkthrough of the method, not a specific client result. We report real numbers only when they are real. The example copy is invented, and no reply rate is attached to it.
-
1Week 0 · Design
Two candidates, checked
A benchmark of how teams their size staff support, against a teardown of their own help center. Different containers, and no single objection kills both.
-
2Week 0 · Freeze
Six lines and a rule
One list build, randomized the day it is made. Three steps, four mailboxes, four whole weeks. Then the kill rule, with the other founder's name on it.
-
3Weeks 1 to 4 · Send
Candidate one, and nothing else
Replies get sorted the day they arrive. No edits, no second candidate alongside, no dashboard check that could turn into a Tuesday decision.
-
4Week 5 · Close
A decision, then seven lines
The rule fires or it does not, on the date either way. Seven lines get written, and candidate two is designed from the objection that came back most.
Five ways an offer test lies to you
Four of these cost you a month. The fifth costs you the only audience you had.
The candidates differ in wording, so a month of sending returns what one minute of the same-objection check would have. The most common failure, and the cheapest to avoid.
Adding an arm feels like learning faster. It splits the same volume three ways and produces three unreadable runs. A third candidate goes next, never alongside.
Day ten, four replies in, the subject line gets edited. The sends before the edit now belong to nothing, and you will not notice until you try to write the log.
Candidate A ran in July into a fresh list, candidate B in September into what was left. Whatever separated them, the offer was one of several differences.
Every test spends the same finite segment
A seed segment is small, and each candidate spends part of it. Run three back to back and everyone in it has heard from you three times in a quarter, with the later ones landing on people who already ignored you. The audience runs out long before significance does.
Questions founders ask
How do you test a cold email offer?
How many emails do you need to test an offer?
Should I test two offers at the same time or one after the other?
How do I know if two offers are different enough to compare?
What is a kill rule in outbound?
What if the offer test is not conclusive?
Co-founder of Real Good GTM. He has been the first business hire and Chief of Staff at seed-stage B2B startups, building outbound pipeline before any playbook existed. This post comes from running offer tests at seed volume, where the afternoon spent choosing two candidates decides more than the four weeks spent sending them.
Connect on LinkedInThe three pages around this one
What an offer is, how the experiment underneath it is designed, and the mirror-image test that moves the audience instead.
The cold email offer
What an offer actually is, its four parts, six shapes that work, and what a test costs at seed volume.
Read the guideCold email A/B testing
The unit you randomize, the confounds outbound has that the web does not, and the pre-send card.
Read the guideICP slice experiments
The mirror image: hold the offer still and move the audience, one narrow slice at a time.
See the playWant the two of us on your offer before you spend a month on it?
Book a fit check. We'll look at the offers you are choosing between, tell you whether a test between them could resolve anything at your volume, and say so straight if the answer is no.
Book a Fit CheckNo hard sell. No fake numbers. Real good work speaks for itself.