Turning outbound into GTM learning
Outbound's job is booked meetings and qualified pipeline, and nothing on this page changes that. While it does that job it also produces the cheapest market research you will ever run, and most teams delete it. This guide is what that evidence can and cannot tell you, and what to write down.
By Kshitij Maheshwari, co-founder · Updated August 2026 · 21 min read
Written by operators who run this motion for seed-stage teams. The arithmetic below is ours and it is reproducible.
What outbound actually measures
A campaign produces four different things, and only two of them count as outbound market feedback.
Outbound market feedback is what a live campaign tells you about your market while it is trying to book meetings: who answers, what they say back, and which of your assumptions nobody bothers to argue with.
Four outputs, one spend, and they do not arrive on the same clock:
How many contacts you could find, how many bounced, how many titles were wrong. Evidence about your list, not about your market.
Objections, misroutes, the incumbent they name, the vocabulary they use for the problem. Fast, specific, and the part almost everyone throws away.
Reply rate per segment. One number, arriving slowly, carrying a far wider margin of error than the dashboard printing it implies.
The thing you are paying for, and the smallest sample on the page. Meetings pay the bills and teach you the least.
Outbound produces two kinds of evidence on two clocks. The words arrive in week one and you can read them. The rate arrives months later and usually still cannot settle the argument you wanted it for.
If you have not written down who you are selling to yet, start with our guide to defining an ideal customer profile. This page is what happens to that definition once it meets a live campaign.
The evidence ladder: what each volume buys you
Four rungs, ordered by what each one costs in sends, so you can look at what you have already sent and know which conclusions you have actually bought.
-
1Day one · Free
Your data is true, or it is not
Bounces, catch-alls, contacts you could not find, titles that were wrong on inspection. Costs nothing, arrives immediately, and tells you about your list.
-
2~200 contacts · Week one
The words
A handful of replies. Not a rate you can defend, and completely enough to read: objections, misroutes, incumbents, vocabulary. Almost all the early value sits here.
-
3~1,000 contacts · Month two
A direction
Your rate is now known to roughly a point either way. Enough to say a segment looks plausibly better and move volume toward it. Not enough to publish.
-
4~2,500 per segment
A ranking you can defend
What it takes to separate a 3% rate from a 4.5% one. Most seed companies never buy this, and the fix is to stop making rung four claims.
- 1 Read replies from day one. Look at rates from month two.
- 2 Two hundred contacts buys words, never a rate you can defend.
- 3 A rate without an interval is a claim without evidence.
- 4 Know which rung you are on before you make the claim.
The motion that runs this ladder is our play on ICP slice experiments, which sizes its batches as "equal, sized to read". The rungs above are what "sized to read" means in contacts.
Rung one: your data is true, or it is not
The most common false conclusion in outbound is a coverage problem wearing the costume of a market verdict.
Read a list problem as a verdict
Nobody in mid-market logistics replied. That segment is not our market.
- ✕Four in ten had no findable address
- ✕Bounced mail never reached a human
- ✕Half the titles were wrong on inspection
Report coverage before verdicts
We reached 61% of that segment. Of the ones who received it, four wrote back.
- ✓No-find rate logged per segment
- ✓Bounces and catch-alls counted first
- ✓The verdict waits for the coverage
A 40% no-find rate is a real finding, and what it tells you is that this segment is hard to reach with the data you bought. That is a sourcing decision, not a market decision, and the two get confused constantly.
Rung two: read the replies, not the rate
Six replies is not a rate, and six replies is plenty of evidence, as long as you read them as sentences instead of counting them.
What they push back on
The same objection three times in your first ten replies is a finding. It tells you what the message has to answer before it earns a meeting, and it costs nothing you were not already spending.
Wrong person, every time
"Talk to ops, not me" repeated across a segment is a targeting correction you can make this week. Founders file it as rejection. It is a map, and it is the cheapest one you will ever be handed.
What they already run
Every named competitor and every in-house workaround is a line in your positioning. Three replies naming the same tool is worth more than a month of guessing at where your category boundary sits.
The words they use
They almost never use your words for the problem. Copy theirs into the next sequence verbatim. It is the cheapest message improvement a seed team gets, and it needs no extra volume.
What nobody argues with
A claim nobody pushes back on is either obviously true or completely invisible to them. Worth finding out which, because the assumptions that never get contested are the ones that quietly stay wrong.
Rungs three and four: when a rate becomes a finding
A reply rate is an estimate with a range around it, and at seed volumes that range is wide enough to hide the entire difference you are arguing about.
| Contacts in one segment | Replies at a true 3% | Observed rate | 95% interval | Widest to narrowest |
|---|---|---|---|---|
| 150 | 4 | 2.7% | 1.0% to 6.7% | 6.4x |
| 200 | 6 | 3.0% | 1.4% to 6.4% | 4.6x |
| 500 | 15 | 3.0% | 1.8% to 4.9% | 2.7x |
| 1,000 | 30 | 3.0% | 2.1% to 4.3% | 2.0x |
| 2,000 | 60 | 3.0% | 2.3% to 3.8% | 1.6x |
| 5,000 | 150 | 3.0% | 2.6% to 3.5% | 1.4x |
Read the top row twice. One reply out of 150 contacts is 0.67%, and its 95% interval runs from 0.12% to 3.68%: a segment that looks dead and a segment humming along at three and a half percent are the same observation.
Two segments at 200 contacts each, one at 5% and one at 2.5%, is a clean doubling, and chance alone would produce a gap that size roughly three times in ten. To separate 3% from 4.5% you need about 2,500 contacts per segment. A doubling from 3% to 6% needs about 750 each. Those are the prices.
The argument I have most often is about which of two adjacent segments is winning. It is almost always the wrong argument. Being wrong there costs a quarter of misallocated volume. Being wrong about whether the offer lands at all costs the year.
Check the denominator before you compare anything
Reply rate over contacts, over emails sent, over delivered mail and over opens are four different numbers sharing one name, and the gap between them is roughly the length of your sequence.
| The denominator | What sits on the bottom | The same campaign reads | Why it moves |
|---|---|---|---|
| Per contact | 1,000 people contacted | 3.0% | One person, counted once. |
| Per email sent | 4,000 emails across four steps | 0.75% | Every follow-up dilutes the rate. |
| Per delivered email | 3,800 after a 5% bounce | 0.79% | Bounces come off the bottom. |
| Per opened email | Not reliably countable | Unusable | Pixel tracking is unreliable and costs deliverability. |
Same campaign, same 30 replies, and the sequence length makes it look four times better or worse depending on which line you quote. Now look at what the platforms publish.
Smartlead's benchmark for the first half of 2026, built from more than 850 million emails sent through its own platform, reports that the median sender earns a reply from about 0.74% of contacts, roughly one reply per 135 people contacted. The top 10% clear 2.63%. Its appendix defines a reply against delivered emails; the figures above are from its per-contact charts instead. Checked August 2026.
Instantly's Cold Email Benchmark Report 2026, covering 1 January to 18 December 2025, puts the platform-wide average at 3.43%, and its methodology note defines that as all replies received, follow-ups included, divided by total emails sent. Checked August 2026.
Both disclose their method, which puts them ahead of most of what you will find, and they still are not comparable: one is a median sender measured per person, the other an aggregate across every email on a platform. So before you compare your campaign to a benchmark, or last month to this month, write one sentence: what is on the bottom of this fraction.
Our own signal-based selling guide carries an 8.5% figure from a 2019 study of 12 million outreach emails and says the sample skews to link building. That number and a 0.74% number can both be true.
The ninety-seven percent you cannot read
Almost everyone you contacted said nothing, and silence is unclassified data rather than a negative answer.
Silence is not a verdict on your market
Reading non-response as rejection is the most common and least supported conclusion in outbound. You will kill a live segment because your mail never arrived, and nothing in your dashboard will ever tell you that is what happened.
Three different populations are hiding inside one silence, and only the first two are fixable:
Bounced, filtered, or sent to an address that stopped working last year. Measurable, and the first thing to fix.
Landed in a tab nobody opens, or under three hundred other emails. Partly fixable through timing and sender reputation.
Read and ignored. The only group that is genuinely market feedback, and the one you can never separate from the other two.
The wider category of evidence that looks solid and is not is covered in when signals mislead. Silence is the largest example of it and the one nobody counts.
Want pipeline built and the evidence written down as it runs?
Book a Fit CheckThe lineage: customer development with a send button
Customer development is Steve Blank's method, set out in The Four Steps to the Epiphany (2005): write down what you believe about your customers, then go outside and test whether they behave the way your model says.
Parts one, five and seven of the memo below are that same idea, aimed at a sequence: write the belief down before you send, mark it confirmed or refuted once the replies are in, and say plainly what you still cannot conclude.
Write the assumption before you send
If you did not write down what you expected, whatever happens will confirm what you already believed, because you will rebuild the expectation afterwards to fit.
Write a belief that cannot lose
We believe mid-market logistics is a strong segment for us.
- ✕No outcome could contradict it
- ✕Names no person and no behavior
- ✕Survives any result you get
Write a belief that can lose
We believe ops leads at Series A logistics firms will name manual carrier reconciliation unprompted.
- ✓One line per segment, dated
- ✓Written before the first send
- ✓Replies either name it or they do not
One falsifiable line per slice, dated, before the first send. That single habit is what separates a campaign that teaches you something from a campaign you narrate afterwards.
The GTM Learning Memo, part by part
One document per ICP, seven parts, kept live rather than delivered: this is the structure you can copy tomorrow morning.
What we believed going in
The dated assumption list, one falsifiable line per slice. Written before the first send, because an expectation reconstructed afterwards always turns out to have been right all along.
What we actually sent
Slices, volumes, offers, signals and the denominator, stated plainly. Skip this part and nobody can interpret the rest of the memo three months from now, including the person who wrote it.
What came back in words
Verbatim replies grouped by theme: objections, misroutes, incumbents, vocabulary. In month one this part is longer than the numbers part, and that is correct rather than a gap.
What came back in numbers
Rate per slice with its interval, never a bare percentage, plus the rung of evidence it sits on. A number without an interval is a claim without evidence attached to it.
What we now believe
Each assumption marked confirmed, refuted or still open, with the rung of evidence that earned the mark. Most of them stay open for longer than anyone would like.
What that changes next
Targeting, offer, signal or channel, each with an owner and a date. A learning that changes nothing was not a learning, it was a paragraph in a status update.
What we are not concluding
The questions this volume cannot answer yet, and roughly what volume each would need. Nobody writes this part. It is the one that makes the other six believable.
An illustrative memo, filled in
An illustrative walkthrough of the method, not a specific client result. We report real numbers only when they are real.
Real Good GTM has no published case studies, by choice. The cards below show the shape of a memo after one cycle, written as judgment rather than measurement. Every figure it might carry is left out on purpose.
What we believed
Three slices of one ICP, each with a written line that could lose. The sharpest of them: ops leads at Series A logistics firms will name manual carrier reconciliation without being prompted.
What we sent
Equal batches per slice, one variable changed at a time, the same four-step sequence everywhere, and the denominator written at the top of the sheet: replies per contact, not per email.
What came back in words
Two slices named a different problem than the one in the email. One kept routing us to finance. Nobody argued with the premise, which is its own finding and not a comfortable one.
What changes, and what does not
The routing correction is safe to act on this week. Which slice replies best is not: those batches sit at rung two, and a ranking needs rung four. That line stays in part seven until the volume exists.
Part seven: what you are deliberately not concluding
The list of questions your volume cannot answer yet is the hardest part to write and the only one nobody else publishes.
Documenting what you learned is not the differentiator. Anyone can send a monthly summary. The part almost nobody writes is the list of things the data does not yet support, and that is the part that makes the other six believable.
Every open question gets a rough price next to it, and the price decides what you do about it:
Which slice really replies best. Roughly another 1,500 contacts each, and the memo says so rather than guessing in the meantime.
Whether the buyer exists but never reads cold email. No amount of extra sending settles that one, at any price.
Why they chose the incumbent. That comes from a call, not a sequence, and the memo should say which one it needs.
The first time I wrote a not-concluding section I expected it to read as weakness. It did the opposite. It is the fastest way I know to stop a founder betting a quarter on a number that cannot hold the weight.
Running this as a two-person team
One memo per ICP, one owner, and three fixed appointments: a weekly half hour, a monthly rate read, and a quarterly re-read of part one.
One owner reads the week's replies and tags them as they go, never retroactively: objection, misroute, incumbent, vocabulary. A doc and a spreadsheet is enough, and it stays enough longer than most people expect.
Update the rate and the interval per slice, then ask one question of each: has this narrowed enough to change a decision. Most months the honest answer is no, and writing that down is the point.
Re-read part one and count how many beliefs survived. Beliefs that never got tested at all are more common than beliefs that turned out wrong, and they are the ones quietly setting your roadmap.
One thing about this loop has genuinely changed since 2024, and one has not:
Sorting a few hundred replies into themes is minutes of work now rather than an afternoon, which makes rung two cheap enough to run properly.
None of it adds sample size. A model summarizing six replies produces a confident paragraph built on six replies.
Buy tooling when the tagging backlog outgrows the half hour, not before. When it does, our signal and intent tools guide is the neutral comparison.
The wiring discipline behind the whole loop, from list to inbox to tagged reply, sits in our GTM engineering guide.
- 1 One memo per ICP, one owner, one weekly half hour.
- 2 Tag replies as you read them, never in a catch-up week.
- 3 Classification got faster. The sample size did not move.
- 4 A doc beats a platform nobody on the team opens.
What this looks like across a first ramp, week by week, is in our seed-stage outbound playbook.
Where the common advice is wrong
The prevailing guidance on validating a segment through outbound is roughly an order of magnitude short, and it is short in a specific, checkable way.
"A hundred and fifty to two hundred contacts per segment gives you statistically meaningful signal. Under one percent means the segment is dead."
- ✕Treats a handful of replies as a rate
- ✕Ranks segments on numbers that overlap
- ✕Kills a slice on a single-figure sample
- ✕Reads the number, deletes the replies
"Two hundred contacts is the right volume for reading replies. It is the wrong volume for ranking anything against anything else."
- ✓One reply in 150 could be 0.12% or 3.68%
- ✓A clean doubling at 200 each could easily be noise
- ✓The words at rung two are the actual payoff
- ✓Rank slices at rung four, or do not rank them
The workflow behind that advice is sound, and we run a version of it in ICP slice experiments. It is only the arithmetic underneath that is wrong, and being wrong there is what turns a good process into confident guessing.
What to realistically expect
Expect fast, specific evidence about your message and slow, blunt evidence about your segments, in that order and never the reverse.
Your data holds up or it does not, and you have enough replies to fix who you write to and what you say. Read the table above before you rank anything.
Your rate per slice is real to roughly a point either way. Enough to move volume toward the slice that looks better and defend the move. Still not enough to say one beats another.
A memo with beliefs marked confirmed, refuted and open, and a shorter open list than you started with. Most founders are surprised how many were never tested at all.
If every slice is quiet, that is an offer problem, and no amount of per-segment arithmetic will find it. Uniform silence is the cheapest diagnosis outbound gives you, and it arrives first. The fix lives in our guide to the outbound offer.
Where this becomes a service, said once: alongside the booked meetings, we keep a live GTM learning memo per ICP for the teams we run outbound for, part seven included. The meetings are what you are paying for. The memo is the second output of the same spend, and the day it starts costing extra it has stopped being cheap.
Questions founders ask
How many emails do I need to send before I can trust a segment?
Why do Smartlead and Instantly publish such different reply rates?
My reply rate dropped from 4% to 2%. What changed?
What if nobody replies at all?
Does a positive reply mean they are a good fit?
What actually goes in a GTM learning memo?
How do I actually learn GTM as a founder?
Co-founder of Real Good GTM. He has been the first business hire and Chief of Staff at seed-stage B2B startups, where reading replies was the fastest market research available and the numbers rarely settled anything on their own. The evidence ladder on this page is the version of that discipline we run for the teams we work with.
Connect on LinkedInFrom the memo back to the work
You have the ladder and the document. These three are the work it sits on: the definition you are testing, the motion that tests it, and the thing to fix when everything is quiet.
Defining an ICP
The definition this page tests. How to write an ideal customer profile specific enough that a campaign can prove it wrong.
Read the guideICP slice experiments
The motion that produces the evidence: three to five slices, equal batches, one variable at a time, with kill and expand rules.
Read the playThe outbound offer
What to fix when every slice is quiet at once. Uniform silence is an offer problem, and it is the cheapest diagnosis you get.
Read the guideWant the pipeline built and the evidence kept?
Book a fit check. We'll look at your ICP, the volume you can realistically run, and what that volume can honestly prove. If outbound is the wrong motion for your stage, we'll tell you that too.
Book a Fit CheckNo hard sell. No fake numbers. Real good work speaks for itself.