Skip to content
GTM Guides

Turning outbound into GTM learning

Outbound's job is booked meetings and qualified pipeline, and nothing on this page changes that. While it does that job it also produces the cheapest market research you will ever run, and most teams delete it. This guide is what that evidence can and cannot tell you, and what to write down.

By Kshitij Maheshwari, co-founder · Updated August 2026 · 21 min read

The short answer Five lines, then the detail
What it teaches you
The words in your replies: who really owns the problem, what they already run, and the vocabulary they use for it.
What it cannot
A ranking of two segments you could defend, unless you have sent thousands of messages into each of them.
What volume buys
Roughly 200 contacts to read replies, 1,000 for a direction, 2,500 per segment to defend a ranking.
What you write down
One memo per ICP: what you believed, what you sent, what came back in words, and what that changes.
The part nobody writes
The list of questions your volume cannot answer yet. That section is what makes the other six believable.

Written by operators who run this motion for seed-stage teams. The arithmetic below is ours and it is reproducible.


What outbound actually measures

A campaign produces four different things, and only two of them count as outbound market feedback.

Definition

Outbound market feedback is what a live campaign tells you about your market while it is trying to book meetings: who answers, what they say back, and which of your assumptions nobody bothers to argue with.

Also called the learning layer · glossary entry

Four outputs, one spend, and they do not arrive on the same clock:

Data truth

How many contacts you could find, how many bounced, how many titles were wrong. Evidence about your list, not about your market.

The words

Objections, misroutes, the incumbent they name, the vocabulary they use for the problem. Fast, specific, and the part almost everyone throws away.

The rates

Reply rate per segment. One number, arriving slowly, carrying a far wider margin of error than the dashboard printing it implies.

The meetings

The thing you are paying for, and the smallest sample on the page. Meetings pay the bills and teach you the least.

Outbound produces two kinds of evidence on two clocks. The words arrive in week one and you can read them. The rate arrives months later and usually still cannot settle the argument you wanted it for.

If you have not written down who you are selling to yet, start with our guide to defining an ideal customer profile. This page is what happens to that definition once it meets a live campaign.


The framework

The evidence ladder: what each volume buys you

Four rungs, ordered by what each one costs in sends, so you can look at what you have already sent and know which conclusions you have actually bought.

  1. 1
    Day one · Free

    Your data is true, or it is not

    Bounces, catch-alls, contacts you could not find, titles that were wrong on inspection. Costs nothing, arrives immediately, and tells you about your list.

  2. 2
    ~200 contacts · Week one

    The words

    A handful of replies. Not a rate you can defend, and completely enough to read: objections, misroutes, incumbents, vocabulary. Almost all the early value sits here.

  3. 3
    ~1,000 contacts · Month two

    A direction

    Your rate is now known to roughly a point either way. Enough to say a segment looks plausibly better and move volume toward it. Not enough to publish.

  4. 4
    ~2,500 per segment

    A ranking you can defend

    What it takes to separate a 3% rate from a 4.5% one. Most seed companies never buy this, and the fix is to stop making rung four claims.

Key takeaways
4 points
  • 1 Read replies from day one. Look at rates from month two.
  • 2 Two hundred contacts buys words, never a rate you can defend.
  • 3 A rate without an interval is a claim without evidence.
  • 4 Know which rung you are on before you make the claim.

The motion that runs this ladder is our play on ICP slice experiments, which sizes its batches as "equal, sized to read". The rungs above are what "sized to read" means in contacts.


Rung one: your data is true, or it is not

The most common false conclusion in outbound is a coverage problem wearing the costume of a market verdict.

Don't

Read a list problem as a verdict

Nobody in mid-market logistics replied. That segment is not our market.

  • Four in ten had no findable address
  • Bounced mail never reached a human
  • Half the titles were wrong on inspection
Do

Report coverage before verdicts

We reached 61% of that segment. Of the ones who received it, four wrote back.

  • No-find rate logged per segment
  • Bounces and catch-alls counted first
  • The verdict waits for the coverage

A 40% no-find rate is a real finding, and what it tells you is that this segment is hard to reach with the data you bought. That is a sourcing decision, not a market decision, and the two get confused constantly.


Rung two: read the replies, not the rate

Six replies is not a rate, and six replies is plenty of evidence, as long as you read them as sentences instead of counting them.

What words buy you
1
The objection

What they push back on

The same objection three times in your first ten replies is a finding. It tells you what the message has to answer before it earns a meeting, and it costs nothing you were not already spending.

2
The misroute

Wrong person, every time

"Talk to ops, not me" repeated across a segment is a targeting correction you can make this week. Founders file it as rejection. It is a map, and it is the cheapest one you will ever be handed.

3
The incumbent

What they already run

Every named competitor and every in-house workaround is a line in your positioning. Three replies naming the same tool is worth more than a month of guessing at where your category boundary sits.

4
The vocabulary

The words they use

They almost never use your words for the problem. Copy theirs into the next sequence verbatim. It is the cheapest message improvement a seed team gets, and it needs no extra volume.

5
Uncontested

What nobody argues with

A claim nobody pushes back on is either obviously true or completely invisible to them. Worth finding out which, because the assumptions that never get contested are the ones that quietly stay wrong.


Rungs three and four: when a rate becomes a finding

A reply rate is an estimate with a range around it, and at seed volumes that range is wide enough to hide the entire difference you are arguing about.

Contacts in one segment Replies at a true 3% Observed rate 95% interval Widest to narrowest
150 4 2.7% 1.0% to 6.7% 6.4x
200 6 3.0% 1.4% to 6.4% 4.6x
500 15 3.0% 1.8% to 4.9% 2.7x
1,000 30 3.0% 2.1% to 4.3% 2.0x
2,000 60 3.0% 2.3% to 3.8% 1.6x
5,000 150 3.0% 2.6% to 3.5% 1.4x
Intervals are the range each rate could really sit in, computed by us at the usual 95% confidence; sample sizes are set so a real doubling shows up four times in five and a fluke fools you about once in twenty. Both are reproducible: Evan Miller's public sample size calculator returns the same numbers, and how we check what we publish is in how we evaluate.
What it costs to call a winner

Read the top row twice. One reply out of 150 contacts is 0.67%, and its 95% interval runs from 0.12% to 3.68%: a segment that looks dead and a segment humming along at three and a half percent are the same observation.

Two segments at 200 contacts each, one at 5% and one at 2.5%, is a clean doubling, and chance alone would produce a gap that size roughly three times in ten. To separate 3% from 4.5% you need about 2,500 contacts per segment. A doubling from 3% to 6% needs about 750 each. Those are the prices.

Operator note
Learned the hard way

The argument I have most often is about which of two adjacent segments is winning. It is almost always the wrong argument. Being wrong there costs a quarter of misallocated volume. Being wrong about whether the offer lands at all costs the year.

RB
Rahul Bageria
Co-founder, Real Good GTM

Check the denominator before you compare anything

Reply rate over contacts, over emails sent, over delivered mail and over opens are four different numbers sharing one name, and the gap between them is roughly the length of your sequence.

One campaign, 1,000 contacts, a four-step sequence and 30 replies, read four ways
The denominator What sits on the bottom The same campaign reads Why it moves
Per contact 1,000 people contacted 3.0% One person, counted once.
Per email sent 4,000 emails across four steps 0.75% Every follow-up dilutes the rate.
Per delivered email 3,800 after a 5% bounce 0.79% Bounces come off the bottom.
Per opened email Not reliably countable Unusable Pixel tracking is unreliable and costs deliverability.

Same campaign, same 30 replies, and the sequence length makes it look four times better or worse depending on which line you quote. Now look at what the platforms publish.

Vendor data · per contact

Smartlead's benchmark for the first half of 2026, built from more than 850 million emails sent through its own platform, reports that the median sender earns a reply from about 0.74% of contacts, roughly one reply per 135 people contacted. The top 10% clear 2.63%. Its appendix defines a reply against delivered emails; the figures above are from its per-contact charts instead. Checked August 2026.

Vendor data · per email sent

Instantly's Cold Email Benchmark Report 2026, covering 1 January to 18 December 2025, puts the platform-wide average at 3.43%, and its methodology note defines that as all replies received, follow-ups included, divided by total emails sent. Checked August 2026.

The rule

Both disclose their method, which puts them ahead of most of what you will find, and they still are not comparable: one is a median sender measured per person, the other an aggregate across every email on a platform. So before you compare your campaign to a benchmark, or last month to this month, write one sentence: what is on the bottom of this fraction.

Our own signal-based selling guide carries an 8.5% figure from a 2019 study of 12 million outreach emails and says the sample skews to link building. That number and a 0.74% number can both be true.


A reply is not an interview

An outbound sequence is a pitch, so every reply you get is a reaction to your idea rather than a description of their life.

Bad data, in a nice suit
  • "Sounds interesting, send more info"
  • "We would definitely look at that next year"
  • "Great idea, good luck with it"
  • Any promise about a future budget
Watch-outs
  • !Compliments feel like signal and are not
  • !Hypotheticals about next quarter never land
  • !Repliers are the few willing to answer strangers
  • !A pattern in replies is a hypothesis, not a finding
Where this limit comes from

Rob Fitzpatrick's The Mom Test (2013) argues that the moment you mention your idea the conversation is contaminated, and that compliments, fluff and hypotheticals are the bad data you collect as a result. A cold sequence mentions your idea in the first line, so it starts contaminated. What survives is narrow and useful: what they already run, who owns the problem, and what is happening to them now.


The ninety-seven percent you cannot read

Almost everyone you contacted said nothing, and silence is unclassified data rather than a negative answer.

!
Caution

Silence is not a verdict on your market

Reading non-response as rejection is the most common and least supported conclusion in outbound. You will kill a live segment because your mail never arrived, and nothing in your dashboard will ever tell you that is what happened.

Do this instead
Measure the causes you can see, then shrink the unclassified pile rather than interpreting it.

Three different populations are hiding inside one silence, and only the first two are fixable:

Never arrived

Bounced, filtered, or sent to an address that stopped working last year. Measurable, and the first thing to fix.

Arrived, never seen

Landed in a tab nobody opens, or under three hundred other emails. Partly fixable through timing and sender reputation.

Seen, not moved

Read and ignored. The only group that is genuinely market feedback, and the one you can never separate from the other two.

The wider category of evidence that looks solid and is not is covered in when signals mislead. Silence is the largest example of it and the one nobody counts.

Want pipeline built and the evidence written down as it runs?

Book a Fit Check

The lineage: customer development with a send button

Definition

Customer development is Steve Blank's method, set out in The Four Steps to the Epiphany (2005): write down what you believe about your customers, then go outside and test whether they behave the way your model says.

Parts one, five and seven of the memo below are that same idea, aimed at a sequence: write the belief down before you send, mark it confirmed or refuted once the replies are in, and say plainly what you still cannot conclude.


Write the assumption before you send

If you did not write down what you expected, whatever happens will confirm what you already believed, because you will rebuild the expectation afterwards to fit.

Don't

Write a belief that cannot lose

We believe mid-market logistics is a strong segment for us.

  • No outcome could contradict it
  • Names no person and no behavior
  • Survives any result you get
Do

Write a belief that can lose

We believe ops leads at Series A logistics firms will name manual carrier reconciliation unprompted.

  • One line per segment, dated
  • Written before the first send
  • Replies either name it or they do not

One falsifiable line per slice, dated, before the first send. That single habit is what separates a campaign that teaches you something from a campaign you narrate afterwards.


The GTM Learning Memo, part by part

One document per ICP, seven parts, kept live rather than delivered: this is the structure you can copy tomorrow morning.

The seven parts
Belief

What we believed going in

The dated assumption list, one falsifiable line per slice. Written before the first send, because an expectation reconstructed afterwards always turns out to have been right all along.

Volume

What we actually sent

Slices, volumes, offers, signals and the denominator, stated plainly. Skip this part and nobody can interpret the rest of the memo three months from now, including the person who wrote it.

Words

What came back in words

Verbatim replies grouped by theme: objections, misroutes, incumbents, vocabulary. In month one this part is longer than the numbers part, and that is correct rather than a gap.

Numbers

What came back in numbers

Rate per slice with its interval, never a bare percentage, plus the rung of evidence it sits on. A number without an interval is a claim without evidence attached to it.

Verdict

What we now believe

Each assumption marked confirmed, refuted or still open, with the rung of evidence that earned the mark. Most of them stay open for longer than anyone would like.

Change

What that changes next

Targeting, offer, signal or channel, each with an owner and a date. A learning that changes nothing was not a learning, it was a paragraph in a status update.

Refusal

What we are not concluding

The questions this volume cannot answer yet, and roughly what volume each would need. Nobody writes this part. It is the one that makes the other six believable.


A swappable skeleton

An illustrative memo, filled in

An illustrative walkthrough of the method, not a specific client result. We report real numbers only when they are real.

i
Illustrative, not a client result

Real Good GTM has no published case studies, by choice. The cards below show the shape of a memo after one cycle, written as judgment rather than measurement. Every figure it might carry is left out on purpose.

Part one

What we believed

Three slices of one ICP, each with a written line that could lose. The sharpest of them: ops leads at Series A logistics firms will name manual carrier reconciliation without being prompted.

Part two

What we sent

Equal batches per slice, one variable changed at a time, the same four-step sequence everywhere, and the denominator written at the top of the sheet: replies per contact, not per email.

Part three

What came back in words

Two slices named a different problem than the one in the email. One kept routing us to finance. Nobody argued with the premise, which is its own finding and not a comfortable one.

Parts five to seven

What changes, and what does not

The routing correction is safe to act on this week. Which slice replies best is not: those batches sit at rung two, and a ranking needs rung four. That line stays in part seven until the volume exists.


Part seven: what you are deliberately not concluding

The list of questions your volume cannot answer yet is the hardest part to write and the only one nobody else publishes.

Documenting what you learned is not the differentiator. Anyone can send a monthly summary. The part almost nobody writes is the list of things the data does not yet support, and that is the part that makes the other six believable.

Every open question gets a rough price next to it, and the price decides what you do about it:

Priced in volume

Which slice really replies best. Roughly another 1,500 contacts each, and the memo says so rather than guessing in the meantime.

Priced in a channel

Whether the buyer exists but never reads cold email. No amount of extra sending settles that one, at any price.

Priced in a conversation

Why they chose the incumbent. That comes from a call, not a sequence, and the memo should say which one it needs.

Operator note
Learned the hard way

The first time I wrote a not-concluding section I expected it to read as weakness. It did the opposite. It is the fastest way I know to stop a founder betting a quarter on a number that cannot hold the weight.

KM
Kshitij Maheshwari
Co-founder, Real Good GTM

Run it small

Running this as a two-person team

One memo per ICP, one owner, and three fixed appointments: a weekly half hour, a monthly rate read, and a quarterly re-read of part one.

Weekly, thirty minutes

One owner reads the week's replies and tags them as they go, never retroactively: objection, misroute, incumbent, vocabulary. A doc and a spreadsheet is enough, and it stays enough longer than most people expect.

Monthly

Update the rate and the interval per slice, then ask one question of each: has this narrowed enough to change a decision. Most months the honest answer is no, and writing that down is the point.

Quarterly

Re-read part one and count how many beliefs survived. Beliefs that never got tested at all are more common than beliefs that turned out wrong, and they are the ones quietly setting your roadmap.

One thing about this loop has genuinely changed since 2024, and one has not:

What changed

Sorting a few hundred replies into themes is minutes of work now rather than an afternoon, which makes rung two cheap enough to run properly.

What did not

None of it adds sample size. A model summarizing six replies produces a confident paragraph built on six replies.

Buy tooling when the tagging backlog outgrows the half hour, not before. When it does, our signal and intent tools guide is the neutral comparison.

The wiring discipline behind the whole loop, from list to inbox to tagged reply, sits in our GTM engineering guide.

Key takeaways
4 points
  • 1 One memo per ICP, one owner, one weekly half hour.
  • 2 Tag replies as you read them, never in a catch-up week.
  • 3 Classification got faster. The sample size did not move.
  • 4 A doc beats a platform nobody on the team opens.

What this looks like across a first ramp, week by week, is in our seed-stage outbound playbook.


Pushback

Where the common advice is wrong

The prevailing guidance on validating a segment through outbound is roughly an order of magnitude short, and it is short in a specific, checkable way.

The common advice

"A hundred and fifty to two hundred contacts per segment gives you statistically meaningful signal. Under one percent means the segment is dead."

  • Treats a handful of replies as a rate
  • Ranks segments on numbers that overlap
  • Kills a slice on a single-figure sample
  • Reads the number, deletes the replies
What actually works
The correction

"Two hundred contacts is the right volume for reading replies. It is the wrong volume for ranking anything against anything else."

  • One reply in 150 could be 0.12% or 3.68%
  • A clean doubling at 200 each could easily be noise
  • The words at rung two are the actual payoff
  • Rank slices at rung four, or do not rank them

The workflow behind that advice is sound, and we run a version of it in ICP slice experiments. It is only the arithmetic underneath that is wrong, and being wrong there is what turns a good process into confident guessing.


What to realistically expect

Expect fast, specific evidence about your message and slow, blunt evidence about your segments, in that order and never the reverse.

After 500 contacts

Your data holds up or it does not, and you have enough replies to fix who you write to and what you say. Read the table above before you rank anything.

After 2,000 contacts

Your rate per slice is real to roughly a point either way. Enough to move volume toward the slice that looks better and defend the move. Still not enough to say one beats another.

After a quarter

A memo with beliefs marked confirmed, refuted and open, and a shorter open list than you started with. Most founders are surprised how many were never tested at all.

The through-line

If every slice is quiet, that is an offer problem, and no amount of per-segment arithmetic will find it. Uniform silence is the cheapest diagnosis outbound gives you, and it arrives first. The fix lives in our guide to the outbound offer.

Where this becomes a service, said once: alongside the booked meetings, we keep a live GTM learning memo per ICP for the teams we run outbound for, part seven included. The meetings are what you are paying for. The memo is the second output of the same spend, and the day it starts costing extra it has stopped being cheap.


FAQ

Questions founders ask

How many emails do I need to send before I can trust a segment?
It depends on what you want to conclude. About 200 contacts gets you enough replies to read: objections, misroutes, the incumbent they name, the words they use. Around 1,000 gets your reply rate known to roughly a point either way. Around 2,500 per segment separates a 3% rate from a 4.5% one at 80% power. Most seed companies never send that much into one segment, and the honest response is to stop making the claims that need it.
Why do Smartlead and Instantly publish such different reply rates?
Different denominators. Smartlead's benchmark for the first half of 2026, from more than 850 million emails sent through its platform, reports the median sender gets a reply from about 0.74% of contacts. Instantly's 2026 report puts its platform-wide average at 3.43%, replies divided by total emails sent. Both disclose their method, and they still are not comparable. How different tools define a reply at all is the wider problem, covered in our guide to outbound metrics.
My reply rate dropped from 4% to 2%. What changed?
Possibly nothing. At the volumes most seed teams send, that swing sits inside the noise: 200 contacts at a true 3% rate produces an observed rate anywhere between about 1.4% and 6.4% with nothing changing underneath. Check deliverability and the denominator first, then check whether the intervals on the two months even separate. They usually do not.
What if nobody replies at all?
Silence is unclassified data, not a verdict. It contains people your mail never reached, people who never saw it, people who were the wrong person, and people who were right and busy. Separate the causes you can measure, bounces, deliverability, wrong contacts, from the one you cannot. If every slice is quiet at once, look at the offer rather than the list.
Does a positive reply mean they are a good fit?
It means they were willing to answer a stranger. That is a real and useful thing and it is not the same thing. Repliers come from the small share of your list most willing to engage at all, so a pattern in your replies is a hypothesis about your market rather than a finding. Write it down that way and it stays useful. Write it down as a finding and it quietly becomes the plan.
What actually goes in a GTM learning memo?
Seven parts, one document per ICP. What we believed going in, written so it could turn out false. What we actually sent, including the denominator. What came back in words, grouped by theme. What came back in numbers, each with its interval. What we now believe, each marked confirmed, refuted or open. What changes next, with an owner and a date. And what we are deliberately not concluding yet, with roughly the volume needed to close each open question.
How do I actually learn GTM as a founder?
By running it and writing down what happened. No course can tell you which objection your buyers raise, which title routes you elsewhere, or which words they use for the problem you solve, because those facts only exist inside your market. Write the belief before you send, read the replies before you read the dashboard, and keep one document per ICP recording what changed your mind. That loop is the whole of this page.
Kshitij Maheshwari, co-founder of Real Good GTM
About the author
Kshitij Maheshwari

Co-founder of Real Good GTM. He has been the first business hire and Chief of Staff at seed-stage B2B startups, where reading replies was the fastest market research available and the numbers rarely settled anything on their own. The evidence ladder on this page is the version of that discipline we run for the teams we work with.

Connect on LinkedIn

Keep going

From the memo back to the work

You have the ladder and the document. These three are the work it sits on: the definition you are testing, the motion that tests it, and the thing to fix when everything is quiet.

Want the pipeline built and the evidence kept?

Book a fit check. We'll look at your ICP, the volume you can realistically run, and what that volume can honestly prove. If outbound is the wrong motion for your stage, we'll tell you that too.

Book a Fit Check

No hard sell. No fake numbers. Real good work speaks for itself.