Skip to content
GTM Guides

Outbound sales metrics: the six you keep

A seed outbound program needs six counts, not fifteen rates. This guide names the six, gives each one a counting card so two people get the same answer, maps the five leaks between them, and shows you how to set a baseline from your own first four weeks.

By Rahul Bageria, co-founder · Updated August 2026 · 21 min read

The short answer Five lines, then the detail
What to keep
Six counts, in stage order: contacts entered, contacts reached, replied, replied with interest, meetings booked, meetings held.
Why counts
A count survives being recounted. A rate is two numbers and a decision somebody made silently.
The artifact
Each count gets a card of five lines: unit, denominator, source of record, exclusions, clock.
The leaks
Five gaps sit between the six, and each one indicts exactly one thing. That is what makes a drop actionable.
The baseline
There is no benchmark you can borrow. Four weeks of your own counts is the band you compare to.

Written by two operators who run this for seed-stage teams, not by a vendor whose dashboard defines the words.


What an outbound metric actually is

A metric is not a thing your tool observes. It is the set of operations that produce a number, so two tools counting differently report two different numbers that share a name.

Definition

An outbound metric is a count plus the rules that made it: the unit, the exclusions, the system you believe, and when the count is fixed. Change one rule and it is a different metric.

Also called an operational definition · full glossary

A number you cannot recount the same way next Tuesday is not a measurement. It is a mood.

P. W. Bridgman argued the same thing about physical length in The Logic of Modern Physics (1927): "we mean by any concept nothing more than a set of operations."


The six counts

The instruments in our cold email deliverability guide answer whether the mail arrived. These six answer what happened after it did.

Six numbers, in stage order
One person

1. Contacts entered

People who entered the campaign, deduped on email address, counted once however many messages they go on to receive.

One person

2. Contacts reached

People for whom at least one message was accepted by the receiving server. Not people who opened anything, and not a count of messages.

One person

3. People who replied

People who sent at least one human reply, whatever it said. A no is a reply. Three replies from one person is still one person.

One person

4. People who replied with interest

The subset of row three that asked for something: time, information, a person, or a price. Where that boundary sits, and why two published rates for it are not comparable, is in our post on positive reply rate.

One event

5. Meetings booked

Events on a calendar with a confirmed attendee. The unit changes here, from a person to an event, and that is the first place a list of numbers usually goes wrong.

One event

6. Meetings held

Events that happened, with a human on the other side. Read from the calendar, never from the tool that booked it.

Everything else on your dashboard is an input to one of these six, an instrument that belongs to deliverability, or decoration.

When you do need a rate, which fraction to compare and how to convert a per-contact one into a per-email one is worked through in our seed stage outbound playbook.


The artifact

The counting card

Every count gets one card, five lines long. Write them before the first send, because a card written after a result exists is a card written to protect the result.

Five lines per number
1
The unit

Person, message, thread or event

Pick one and never mix. Apollo's sequence documentation defines its reply metric as the percentage of contacts who replied "out of the total number of emails sent", which is two units in one fraction.

2
The denominator

What sits underneath, if anything

If the number is a count, write "none" and move on. That is a complete answer, and it is most of the reason to keep counts.

3
Source of record

One system, named

Whose version you believe when two disagree. Not "the CRM and the sequencer". One of them, written down, for each of the six.

4
Exclusions

The list of what does not count

Out-of-office. Auto-responders. A colleague on the same domain. A reply to your own calendar invite. Your own test sends. Anything you did not think of joins the list the first time it happens.

5
The clock

When the number is fixed

A meeting booked on the 31st for the 4th belongs to one month or the other. Decide once, and a chart stops changing every time you rebuild it.

The test that makes a card real: hand it to the other founder and have them recount last week from the raw records. If the two numbers differ, the card is wrong, not the counter.

Operator note
Learned the hard way

The argument two founders have most is not about the number, it is about who counted it. One of you exported from the sequencer, the other read the CRM, and the two disagree by a third. Nobody is wrong. The card ends that argument permanently and takes an afternoon to write.

RB
Rahul Bageria
Co-founder, Real Good GTM

The evidence

Six tools disagree about what a reply is

All six answers are published, in the vendors' own help centers, one click from the product. This is the numerator and the exclusion list, not the denominator.

The tool What one reply is What its own documentation adds
Instantly A lead who replied to at least one email in the selected range. Automatic replies are included by default, and the switch that excludes them stays off "for all campaigns on your current device", not on the account.
lemlist A lead who responded to any step of the campaign. "Auto-replies and out-of-office messages aren't tracked", and reply tracking is a per-campaign toggle that logs nothing at all when it is off.
Smartlead A reply event, not a person. "Each reply is counted; multiple replies from the same lead count individually", so one talkative prospect lands three times.
Apollo A contact who replied to at least one email in the sequence. It "automatically categorizes any reply, including to meeting invites, as a reply", including threads it did not start. Booking the meeting moves the number.
Saleshandy A conversation, tracked and filed per sequence. "Replies from anyone with the same domain as your saved prospect will also be tracked", and a forwarded or CC'd address "will be created as a new prospect".
HubSpot "The percentage of contacts who replied at least once to any email sent from this sequence." Contacts are deduplicated by email address, so one person is one record however many rows your import had.

Quoted from each vendor's own help center, read 16 August 2026; the pages run from January 2025 to August 2026, and Smartlead's stamps none of its articles. lemlist does detect an out-of-office and reschedule around it. It just never logs one as a reply, the exact opposite of Instantly's default.


The label is not the formula

One row on one widely used dashboard describes three different fractions at once. The vendor's own research is on the right side of this, which is what makes the row interesting rather than damning.

What the label says

"Reply Rate: Percentage of recipients who replied."

  • Reads as replies over recipients
  • Reads as one reply counted per person
  • Reads as everybody you contacted
  • Reads as comparable to a per-contact figure
What the formula computes

"(Total replies ÷ unique opens) × 100"

  • The numerator counts reply events, not people
  • The row above it: replies from one lead count individually
  • The denominator is openers, so non-openers are not in it
  • Replies over recipients is a fraction nobody computed

Both cells are Smartlead's, side by side in one table. Its own research says the opposite of its dashboard. The State of Cold Email report, on more than 850 million emails sent through the platform from January to June 2026, is blunt about open rate:

The vendor's own report, page 19

"We deliberately do not use it. Apple Mail loads email images automatically and triggers open tracking before anyone reads the message, so reported open rates of 60 to 70 percent are largely automatic. Only a reply shows genuine interest."

We agree with every word of that, and the agreement is the point. A product and the research published about it drift apart when a definition ends up living in a table cell nobody owns.

It leaves one question worth asking your own vendor rather than assuming the answer: if a reply rate divides by unique opens, what does it report for a team that has turned open tracking off?

Label and formula quoted from Smartlead's main dashboard analytics article, checked August 2026; that help center publishes no dates. The report quote sits in appendix A, page 19 of 22, and carries no publication date either, so its dataset period is the stamp.


Check your own help center before you trust a number

Twenty minutes, once per tool. Every line below is answered somewhere in your vendor's own documentation, and none of it is linked from the dashboard.

Before you trust the number

5 checks

  • Find the page that prints the arithmetic

    Search the help center for the metric name plus "calculation". If nothing shows a formula, that is itself the answer.

  • Read the label and the formula separately

    Write both down. If they describe different fractions, the formula is what your chart is made of.

  • Establish whether it counts people or events

    Reply twice from one test address and watch whether the number moves once or twice.

  • Find the auto-reply setting, and where it lives

    Some are per campaign, some per account, one is per device. Check it on the machine you read the number from.

  • Note the date on the page, if there is one

    Undated documentation is not wrong. It means the only stamp you own is the day you read it.

Whatever you find goes on the card, under exclusions. Re-run it quarterly: a definition that changes under a trend line looks exactly like a market that changed.


Where each number comes from, and what breaks it

Every count is produced by one system, and every one of them has a failure that looks like a change in the market rather than a change in the plumbing.

The count Where it physically comes from What quietly breaks it
Contacts entered The list you uploaded, held in the sequencer. A re-upload creates a second record for the same person, and both of them get counted.
Contacts reached The sequencer's delivery log. A rejected message comes back with a code saying whether the address is dead or the mailbox was busy. Hard and soft are each vendor's own reading of that code, so two tools count one list differently.
People who replied Your mailbox, read by the sequencer. Forwarding rules and stripped headers break reply matching, and auto-responders land in the count in one tool and never in another.
Replied with interest Whoever classifies the reply, human or model. Outcomes filed per sequence rather than per person let the same human be two things at once in your reporting.
Meetings booked The calendar, or the booking tool that wrote to it. Reporting time zones and month boundaries. A meeting booked on the last day for the fourth of the next belongs wherever your clock says.
Meetings held The calendar, after the fact, marked by hand. Nothing marks a no-show automatically. If booked and held look nearly identical, held is coming from the wrong system.

Read the right-hand column as a list of things that move a number without anything happening in your market.


Every number needs a counter-number

An indicator is not neutral. It steers you toward whatever it monitors, so each count needs the number that stops you gaming it.

Andrew Grove put it plainly in High Output Management (1983): "Indicators tend to direct your attention toward what they are monitoring. It is like riding a bicycle: you will probably steer it where you are looking."

His fix is to pair them, "so that together both effect and counter-effect are measured." Read the columns below as six pairs, row by row.

The count you keep
  • Contacts entered
  • Contacts reached
  • People who replied
  • People who replied with interest
  • Meetings booked
  • Meetings held
Read it next to this one
  • !Contacts you could not verify or disqualified
  • !Hard bounces, read in the deliverability instruments
  • !Opt-outs
  • !Interested replies that never reached a calendar
  • !Meetings held
  • !Meetings that turned out to be off ICP
The through-line

Read the right-hand column on its own and it is a list of every way an outbound team has ever flattered itself: padding the list, buying delivery with a dirty one, writing to provoke, marking anything polite as interested, booking whoever says yes.

A single number is an instruction to move it. A pair is an instruction to move one without wrecking the other.

Want the counting set up properly before the first send?

Book a Fit Check

Sends do not belong on the list

A send count is the easiest number on any dashboard to move, and the only one a person can change with a slider. That is precisely why it is not one of the six.

!
Caution

A send count on a list becomes a target

The moment a number becomes the target, it stops telling you anything. Put sends on the list you read to decide things, and the only decision left on it is to send more.

Grove settles it: "you measure a salesman by the orders he gets (output), not by the calls he makes (activity)." A send is an activity, and so is a touch.

Do this instead
Keep sends on the operating checklist that says the machine ran this week. Not on the list you read to decide anything.

Opens, in one sentence

The hand-off

An open is a request for a hidden image, and a large share of those requests are now machines rather than people. It is not one of the six, and it is not a success metric.

The reading that does survive, and why it is narrower than anyone wants, is in our post on cold email subject lines. This guide has nothing to add.


The five leaks, and what each one indicts

Five gaps sit between the six counts. Each one indicts exactly one thing, which is what turns a worrying chart into a decision about what to work on next.

Five gaps, five verdicts
1
Data and setup

Entered to reached: the mail did not arrive

Verification, the addresses themselves, the sending setup. Nothing about the message is on trial here, and this is the one leak that belongs to a different guide.

2
List or offer

Reached to replied: it arrived and nobody cared

Two candidates, and the tell is old. Two different angles failing in the same way means the offer, not the list.

3
The ask

Replied to interested: they answered, none of it warm

Plenty of replies and no interest in any of them is a problem with what you asked for, not with who you asked.

4
The handoff

Interested to booked: somebody said yes, then nothing

The cheapest leak to fix and the last one anyone looks at. Somebody has to classify and answer every reply the hour it lands, which is the whole of reply automation.

5
The confirmation

Booked to held: the slot came and went

Reminders, the gap between the booking and the slot, and whether the meeting was ever real. Read the next section before you read this leak.

The chart looks the same in all five cases. Only the gap tells you which one you are in.


Booked is not held

A booked meeting has three fates, not two, so held rate and no-show rate are not complements of each other.

6,428
Scheduled meetings, one week

is the sample behind the only no-show figure on this page.

As of August 2026 Inbound scheduling

Of those 6,428, 4,895 were completed and 419 were no-shows. Subtract both and 1,114 are left, 17.3% of the total, in neither bucket: canceled, rescheduled, or still in the future when the count was taken.

RevenueHero, no-show benchmark report, published 13 December 2024. Meetings booked through its own scheduling product, 15 industries, one week. Its own note: "we found some meetings that are still scheduled and set to happen in the future."

Two caveats belong in the same breath as that figure: it is inbound scheduling traffic rather than outbound-booked meetings, and it is one week of it. Take the third bucket from it, never a no-show rate for yourself.


"A meeting" means four different things

This is not a cold email quirk. The Bridge Group's 2025 sales development report publishes four separate quotas for one job, because the word covers four different bars. Two of them make the point.

The bar: it happened
16.0
Meetings a month, one rep

An introductory meeting, where showing up is the whole test.

The bar: it qualified
9.0
Meetings a month, one rep

A fully qualified meeting, which cleared a definition somebody wrote down first.

The move

Same job, same month, and the target moves with the word the company uses.

Write on the counting card which bar your meeting clears: attendance, or a test somebody wrote down first. Then never compare your count to anyone whose bar you have not read.

The Bridge Group, Sales Development Models, Metrics and Compensation, 10th edition, published February 2025. Both figures are medians across mostly North American B2B software teams.


What to read weekly, and what to read monthly

Most of these numbers cannot legitimately move week to week at seed volume. A cadence is how you stop reading noise as news.

  1. 1

    Daily: two alarms, and no rates at all

    A mailbox that stopped sending, and a bounce spike. Nothing among the six moves meaningfully in a day, so nothing among the six is a daily number.

  2. 2

    Weekly: opt-outs, and the replies as sentences

    Read what people wrote, not how many wrote. A weekly reply rate at seed volume is noise with a decimal point on it.

  3. 3

    Monthly: the six counts, stage to stage

    One segment at a time, never a single number alone. Once a quarter, re-read the counting card against the tools' current documentation, because vendors change definitions without sending anyone a memo.


Instead of a benchmark

Set your own baseline

A published benchmark was counted with a different set of operations, on a population you are not in, starting from a different answer to what a reply is. It is worth a direction, never a level.

  1. 1

    Write the counting card before the first send

    Not after. A baseline built from numbers whose rules changed halfway through is not a baseline, it is two half-measurements stacked.

  2. 2

    Run four weeks without judging anything

    Read the replies as sentences the whole time. Four weeks is long enough to own counts, and nowhere near long enough to own a rate you could defend.

  3. 3

    Take the four-week totals as counts

    Six numbers, written down with the date and the version of the counting card that produced them. That pair is the baseline.

  4. 4

    Compare to the band, not to a number

    Next month is better or worse than your own range. Re-baseline when you change the operations, never when the number moves.

Operator note
What it feels like

The first month's baseline always looks too low to say out loud, and that feeling is correct rather than useful. It is the only honest number you own, and it beats a benchmark from a population you are not in.

KM
Kshitij Maheshwari
Co-founder, Real Good GTM

What to stop tracking

Nobody removes a metric, because removing one feels like admitting it should never have been added. Set the cap at six and enforce it by deleting.

Don't

Keep the fifteen-column dashboard

Sends / Opens / Open rate / Clicks / Click rate / Replies / Reply rate / Health score

  • A composite score hides which part moved
  • Open and click rates are instrument readings
  • Per-step rates split three replies five ways
Do

Keep six columns and delete the rest

Entered / Reached / Replied / Interested / Booked / Held

  • Every column is a count you can recount
  • Fractions get formed when you need one
  • Sends live on the operating checklist

Write down what you decided not to count, and why. Somebody will propose adding it back with great confidence, and the note is how you have that argument once.


How we would run it

Running the six counts as a two-person team

An illustrative walkthrough of the method, not a specific client result. We report real numbers only when they are real.

  1. 1
    Before the first send · The sheet

    Six columns and one card tab

    One row per week, six columns for the counts, one tab holding the counting card. The card names the source of record for each column.

  2. 2
    Week one · The exclusions

    The first edge cases arrive

    An out-of-office, and a reply from a colleague on the prospect's domain. Both go on the exclusions line the day they turn up, not later.

  3. 3
    Every Friday · The words

    Ten minutes, nothing divided

    Opt-outs, and the replies read as sentences. The counts go in the row. Nothing gets divided by anything this week.

  4. 4
    End of month · The read

    Stage to stage, one segment

    The five gaps get read in order, one segment at a time, against your own prior months. The card version goes in the row with the counts.

Operator note
Learned the hard way

Every tool you add makes this worse before it makes it better. A new sequencer, a new CRM, a new booking link: each one arrives with its own version of every number. Decide the source of record for all six on the day the tool lands, not the month you notice they disagree.

RB
Rahul Bageria
Co-founder, Real Good GTM

Pushback

Where the common advice is wrong

The pages that rank for this query do the same three things: name eight to twenty metrics, describe each in one sentence, and attach a benchmark range with no denominator behind it.

The common advice

"Track these fifteen KPIs, and here is the benchmark range for each one."

  • More metrics is more visibility
  • A range tells whether your number is good
  • The metric name is the metric
  • Rates first, and counts never
What actually works

"Keep six counts, write down how each one is counted, and compare next month to your own last month."

  • A list past six has lost the argument it is making
  • A borrowed range was counted with other operations
  • The name is the label; the formula is the metric
  • Counts survive a recount, and rates do not

What to realistically expect

Six counts and a card tell you where the motion leaks. They do not tell you whether the offer is right, and they do not produce volume.

What it gives you

A number that means the same thing in March as it did in January, a leak you can name instead of a month you can worry about, and an argument that ends because both of you counted the same way.

What it does not

It will not make the list bigger, and it will not write a better offer. A measurement system measures. It does not produce companies worth writing to, and it never has.


What to take away

Fewer numbers, written down once, read on a schedule. Everything else on this page is detail underneath those five lines.

Key takeaways
5 points
  • 1 Six counts, not fifteen rates. A count survives being recounted.
  • 2 Write the counting card before the first send, never after.
  • 3 Read the formula under the label. They are not always the same fraction.
  • 4 Every count needs the number that stops you gaming it.
  • 5 Your baseline is four weeks of your own counts, not a benchmark.

FAQ

Questions founders ask

What outbound sales metrics should I actually track?
Six counts, in stage order: contacts entered, contacts reached, people who replied, people who replied with interest, meetings booked, meetings held. Everything else on a dashboard is an input to one of those or decoration. The number of metrics is not the hard part. Writing down how each one is counted, so that two people get the same answer, is.
What actually counts as a reply?
Not the same thing in any two tools, and all of it is documented. lemlist does not track out-of-office messages as replies, while Instantly includes automatic replies by default. Apollo counts a reply to your own calendar invite, and Saleshandy one from anyone on the prospect's domain. Write your own exclusion list before you trust the number.
Do out-of-office replies count as replies?
Not consistently. lemlist's documentation says auto-replies and out-of-office messages are not tracked, though it does detect them and reschedule the follow-up. Instantly's says automatic replies are included by default, and the setting that excludes them is stored per device rather than per account. Whatever your tool does, write the exclusion on the card.
What is a good reply rate for cold email?
There is no answer that survives the question 'counted how'. Six widely used tools disagree about what event is a reply at all, and the denominators differ by roughly the length of your sequence. Run four weeks, take your own six counts as the baseline, and compare next month to that band instead.
Why do my sequencer and my CRM report different numbers?
Because they count different objects. HubSpot deduplicates contacts by email address, so one person is one record. Sending tools count leads, messages or threads depending on the screen, and at least one creates a new record out of a forwarded address. Pick one system as the source of record for each of the six and write it down.
What is the difference between a booked meeting and a held meeting?
A booked meeting is a calendar event. A held meeting happened. They are not complements, because a booked meeting can also be canceled, rescheduled, or still in the future when you count. In one vendor dataset of 6,428 scheduled meetings, 4,895 completed and 419 were no-shows, which leaves 17.3% in neither bucket.
How do I set a baseline with no benchmark to compare to?
Write the counting card before the first send, run four weeks without judging anything, then take the four-week totals as counts. That is your baseline, and it is a band rather than a point. Re-baseline when you change the operations, not when the number moves.
Rahul Bageria, co-founder of Real Good GTM
About the author
Rahul Bageria

Co-founder of Real Good GTM. He has been the first business hire and Chief of Staff at seed-stage B2B startups, building outbound pipeline before any playbook existed. He owns the systems and data side of the motion, which is where this guide comes from: the week somebody asks how it is going and two exports disagree by a third.

Connect on LinkedIn

Keep going

Once the counting is settled

These three pick up where this one stops: what a number can prove, which step-level numbers are worth keeping, and what a meeting costs to produce.

Want the outbound run, and the counting done properly?

Book a fit check. We'll look at your ICP, how you sell today, and what you can honestly measure at your volume. If outbound is not the right motion for your stage, we'll tell you that too.

Book a Fit Check

No hard sell. No fake numbers. Real good work speaks for itself.