Outbound sales metrics: the six you keep
A seed outbound program needs six counts, not fifteen rates. This guide names the six, gives each one a counting card so two people get the same answer, maps the five leaks between them, and shows you how to set a baseline from your own first four weeks.
By Rahul Bageria, co-founder · Updated August 2026 · 21 min read
Written by two operators who run this for seed-stage teams, not by a vendor whose dashboard defines the words.
What an outbound metric actually is
A metric is not a thing your tool observes. It is the set of operations that produce a number, so two tools counting differently report two different numbers that share a name.
An outbound metric is a count plus the rules that made it: the unit, the exclusions, the system you believe, and when the count is fixed. Change one rule and it is a different metric.
A number you cannot recount the same way next Tuesday is not a measurement. It is a mood.
P. W. Bridgman argued the same thing about physical length in The Logic of Modern Physics (1927): "we mean by any concept nothing more than a set of operations."
The six counts
The instruments in our cold email deliverability guide answer whether the mail arrived. These six answer what happened after it did.
1. Contacts entered
People who entered the campaign, deduped on email address, counted once however many messages they go on to receive.
2. Contacts reached
People for whom at least one message was accepted by the receiving server. Not people who opened anything, and not a count of messages.
3. People who replied
People who sent at least one human reply, whatever it said. A no is a reply. Three replies from one person is still one person.
4. People who replied with interest
The subset of row three that asked for something: time, information, a person, or a price. Where that boundary sits, and why two published rates for it are not comparable, is in our post on positive reply rate.
5. Meetings booked
Events on a calendar with a confirmed attendee. The unit changes here, from a person to an event, and that is the first place a list of numbers usually goes wrong.
6. Meetings held
Events that happened, with a human on the other side. Read from the calendar, never from the tool that booked it.
Everything else on your dashboard is an input to one of these six, an instrument that belongs to deliverability, or decoration.
When you do need a rate, which fraction to compare and how to convert a per-contact one into a per-email one is worked through in our seed stage outbound playbook.
The counting card
Every count gets one card, five lines long. Write them before the first send, because a card written after a result exists is a card written to protect the result.
Person, message, thread or event
Pick one and never mix. Apollo's sequence documentation defines its reply metric as the percentage of contacts who replied "out of the total number of emails sent", which is two units in one fraction.
What sits underneath, if anything
If the number is a count, write "none" and move on. That is a complete answer, and it is most of the reason to keep counts.
One system, named
Whose version you believe when two disagree. Not "the CRM and the sequencer". One of them, written down, for each of the six.
The list of what does not count
Out-of-office. Auto-responders. A colleague on the same domain. A reply to your own calendar invite. Your own test sends. Anything you did not think of joins the list the first time it happens.
When the number is fixed
A meeting booked on the 31st for the 4th belongs to one month or the other. Decide once, and a chart stops changing every time you rebuild it.
The test that makes a card real: hand it to the other founder and have them recount last week from the raw records. If the two numbers differ, the card is wrong, not the counter.
The argument two founders have most is not about the number, it is about who counted it. One of you exported from the sequencer, the other read the CRM, and the two disagree by a third. Nobody is wrong. The card ends that argument permanently and takes an afternoon to write.
Six tools disagree about what a reply is
All six answers are published, in the vendors' own help centers, one click from the product. This is the numerator and the exclusion list, not the denominator.
| The tool | What one reply is | What its own documentation adds |
|---|---|---|
| Instantly | A lead who replied to at least one email in the selected range. | Automatic replies are included by default, and the switch that excludes them stays off "for all campaigns on your current device", not on the account. |
| lemlist | A lead who responded to any step of the campaign. | "Auto-replies and out-of-office messages aren't tracked", and reply tracking is a per-campaign toggle that logs nothing at all when it is off. |
| Smartlead | A reply event, not a person. | "Each reply is counted; multiple replies from the same lead count individually", so one talkative prospect lands three times. |
| Apollo | A contact who replied to at least one email in the sequence. | It "automatically categorizes any reply, including to meeting invites, as a reply", including threads it did not start. Booking the meeting moves the number. |
| Saleshandy | A conversation, tracked and filed per sequence. | "Replies from anyone with the same domain as your saved prospect will also be tracked", and a forwarded or CC'd address "will be created as a new prospect". |
| HubSpot | "The percentage of contacts who replied at least once to any email sent from this sequence." | Contacts are deduplicated by email address, so one person is one record however many rows your import had. |
Quoted from each vendor's own help center, read 16 August 2026; the pages run from January 2025 to August 2026, and Smartlead's stamps none of its articles. lemlist does detect an out-of-office and reschedule around it. It just never logs one as a reply, the exact opposite of Instantly's default.
The label is not the formula
One row on one widely used dashboard describes three different fractions at once. The vendor's own research is on the right side of this, which is what makes the row interesting rather than damning.
"Reply Rate: Percentage of recipients who replied."
- ✕Reads as replies over recipients
- ✕Reads as one reply counted per person
- ✕Reads as everybody you contacted
- ✕Reads as comparable to a per-contact figure
"(Total replies ÷ unique opens) × 100"
- ✓The numerator counts reply events, not people
- ✓The row above it: replies from one lead count individually
- ✓The denominator is openers, so non-openers are not in it
- ✓Replies over recipients is a fraction nobody computed
Both cells are Smartlead's, side by side in one table. Its own research says the opposite of its dashboard. The State of Cold Email report, on more than 850 million emails sent through the platform from January to June 2026, is blunt about open rate:
"We deliberately do not use it. Apple Mail loads email images automatically and triggers open tracking before anyone reads the message, so reported open rates of 60 to 70 percent are largely automatic. Only a reply shows genuine interest."
We agree with every word of that, and the agreement is the point. A product and the research published about it drift apart when a definition ends up living in a table cell nobody owns.
It leaves one question worth asking your own vendor rather than assuming the answer: if a reply rate divides by unique opens, what does it report for a team that has turned open tracking off?
Label and formula quoted from Smartlead's main dashboard analytics article, checked August 2026; that help center publishes no dates. The report quote sits in appendix A, page 19 of 22, and carries no publication date either, so its dataset period is the stamp.
Check your own help center before you trust a number
Twenty minutes, once per tool. Every line below is answered somewhere in your vendor's own documentation, and none of it is linked from the dashboard.
Before you trust the number
5 checks
-
Find the page that prints the arithmetic
Search the help center for the metric name plus "calculation". If nothing shows a formula, that is itself the answer.
-
Read the label and the formula separately
Write both down. If they describe different fractions, the formula is what your chart is made of.
-
Establish whether it counts people or events
Reply twice from one test address and watch whether the number moves once or twice.
-
Find the auto-reply setting, and where it lives
Some are per campaign, some per account, one is per device. Check it on the machine you read the number from.
-
Note the date on the page, if there is one
Undated documentation is not wrong. It means the only stamp you own is the day you read it.
Whatever you find goes on the card, under exclusions. Re-run it quarterly: a definition that changes under a trend line looks exactly like a market that changed.
Where each number comes from, and what breaks it
Every count is produced by one system, and every one of them has a failure that looks like a change in the market rather than a change in the plumbing.
| The count | Where it physically comes from | What quietly breaks it |
|---|---|---|
| Contacts entered | The list you uploaded, held in the sequencer. | A re-upload creates a second record for the same person, and both of them get counted. |
| Contacts reached | The sequencer's delivery log. | A rejected message comes back with a code saying whether the address is dead or the mailbox was busy. Hard and soft are each vendor's own reading of that code, so two tools count one list differently. |
| People who replied | Your mailbox, read by the sequencer. | Forwarding rules and stripped headers break reply matching, and auto-responders land in the count in one tool and never in another. |
| Replied with interest | Whoever classifies the reply, human or model. | Outcomes filed per sequence rather than per person let the same human be two things at once in your reporting. |
| Meetings booked | The calendar, or the booking tool that wrote to it. | Reporting time zones and month boundaries. A meeting booked on the last day for the fourth of the next belongs wherever your clock says. |
| Meetings held | The calendar, after the fact, marked by hand. | Nothing marks a no-show automatically. If booked and held look nearly identical, held is coming from the wrong system. |
Read the right-hand column as a list of things that move a number without anything happening in your market.
Every number needs a counter-number
An indicator is not neutral. It steers you toward whatever it monitors, so each count needs the number that stops you gaming it.
Andrew Grove put it plainly in High Output Management (1983): "Indicators tend to direct your attention toward what they are monitoring. It is like riding a bicycle: you will probably steer it where you are looking."
His fix is to pair them, "so that together both effect and counter-effect are measured." Read the columns below as six pairs, row by row.
- ✓Contacts entered
- ✓Contacts reached
- ✓People who replied
- ✓People who replied with interest
- ✓Meetings booked
- ✓Meetings held
- !Contacts you could not verify or disqualified
- !Hard bounces, read in the deliverability instruments
- !Opt-outs
- !Interested replies that never reached a calendar
- !Meetings held
- !Meetings that turned out to be off ICP
Read the right-hand column on its own and it is a list of every way an outbound team has ever flattered itself: padding the list, buying delivery with a dirty one, writing to provoke, marking anything polite as interested, booking whoever says yes.
A single number is an instruction to move it. A pair is an instruction to move one without wrecking the other.
Want the counting set up properly before the first send?
Book a Fit CheckSends do not belong on the list
A send count is the easiest number on any dashboard to move, and the only one a person can change with a slider. That is precisely why it is not one of the six.
A send count on a list becomes a target
The moment a number becomes the target, it stops telling you anything. Put sends on the list you read to decide things, and the only decision left on it is to send more.
Grove settles it: "you measure a salesman by the orders he gets (output), not by the calls he makes (activity)." A send is an activity, and so is a touch.
Opens, in one sentence
An open is a request for a hidden image, and a large share of those requests are now machines rather than people. It is not one of the six, and it is not a success metric.
The reading that does survive, and why it is narrower than anyone wants, is in our post on cold email subject lines. This guide has nothing to add.
The five leaks, and what each one indicts
Five gaps sit between the six counts. Each one indicts exactly one thing, which is what turns a worrying chart into a decision about what to work on next.
Entered to reached: the mail did not arrive
Verification, the addresses themselves, the sending setup. Nothing about the message is on trial here, and this is the one leak that belongs to a different guide.
Reached to replied: it arrived and nobody cared
Two candidates, and the tell is old. Two different angles failing in the same way means the offer, not the list.
Replied to interested: they answered, none of it warm
Plenty of replies and no interest in any of them is a problem with what you asked for, not with who you asked.
Interested to booked: somebody said yes, then nothing
The cheapest leak to fix and the last one anyone looks at. Somebody has to classify and answer every reply the hour it lands, which is the whole of reply automation.
Booked to held: the slot came and went
Reminders, the gap between the booking and the slot, and whether the meeting was ever real. Read the next section before you read this leak.
The chart looks the same in all five cases. Only the gap tells you which one you are in.
Booked is not held
A booked meeting has three fates, not two, so held rate and no-show rate are not complements of each other.
is the sample behind the only no-show figure on this page.
Of those 6,428, 4,895 were completed and 419 were no-shows. Subtract both and 1,114 are left, 17.3% of the total, in neither bucket: canceled, rescheduled, or still in the future when the count was taken.
RevenueHero, no-show benchmark report, published 13 December 2024. Meetings booked through its own scheduling product, 15 industries, one week. Its own note: "we found some meetings that are still scheduled and set to happen in the future."
Two caveats belong in the same breath as that figure: it is inbound scheduling traffic rather than outbound-booked meetings, and it is one week of it. Take the third bucket from it, never a no-show rate for yourself.
"A meeting" means four different things
This is not a cold email quirk. The Bridge Group's 2025 sales development report publishes four separate quotas for one job, because the word covers four different bars. Two of them make the point.
An introductory meeting, where showing up is the whole test.
A fully qualified meeting, which cleared a definition somebody wrote down first.
Same job, same month, and the target moves with the word the company uses.
Write on the counting card which bar your meeting clears: attendance, or a test somebody wrote down first. Then never compare your count to anyone whose bar you have not read.
The Bridge Group, Sales Development Models, Metrics and Compensation, 10th edition, published February 2025. Both figures are medians across mostly North American B2B software teams.
What to read weekly, and what to read monthly
Most of these numbers cannot legitimately move week to week at seed volume. A cadence is how you stop reading noise as news.
-
1
Daily: two alarms, and no rates at all
A mailbox that stopped sending, and a bounce spike. Nothing among the six moves meaningfully in a day, so nothing among the six is a daily number.
-
2
Weekly: opt-outs, and the replies as sentences
Read what people wrote, not how many wrote. A weekly reply rate at seed volume is noise with a decimal point on it.
-
3
Monthly: the six counts, stage to stage
One segment at a time, never a single number alone. Once a quarter, re-read the counting card against the tools' current documentation, because vendors change definitions without sending anyone a memo.
Set your own baseline
A published benchmark was counted with a different set of operations, on a population you are not in, starting from a different answer to what a reply is. It is worth a direction, never a level.
-
1
Write the counting card before the first send
Not after. A baseline built from numbers whose rules changed halfway through is not a baseline, it is two half-measurements stacked.
-
2
Run four weeks without judging anything
Read the replies as sentences the whole time. Four weeks is long enough to own counts, and nowhere near long enough to own a rate you could defend.
-
3
Take the four-week totals as counts
Six numbers, written down with the date and the version of the counting card that produced them. That pair is the baseline.
-
4
Compare to the band, not to a number
Next month is better or worse than your own range. Re-baseline when you change the operations, never when the number moves.
The first month's baseline always looks too low to say out loud, and that feeling is correct rather than useful. It is the only honest number you own, and it beats a benchmark from a population you are not in.
What to stop tracking
Nobody removes a metric, because removing one feels like admitting it should never have been added. Set the cap at six and enforce it by deleting.
Keep the fifteen-column dashboard
Sends / Opens / Open rate / Clicks / Click rate / Replies / Reply rate / Health score
- ✕A composite score hides which part moved
- ✕Open and click rates are instrument readings
- ✕Per-step rates split three replies five ways
Keep six columns and delete the rest
Entered / Reached / Replied / Interested / Booked / Held
- ✓Every column is a count you can recount
- ✓Fractions get formed when you need one
- ✓Sends live on the operating checklist
Write down what you decided not to count, and why. Somebody will propose adding it back with great confidence, and the note is how you have that argument once.
Running the six counts as a two-person team
An illustrative walkthrough of the method, not a specific client result. We report real numbers only when they are real.
-
1Before the first send · The sheet
Six columns and one card tab
One row per week, six columns for the counts, one tab holding the counting card. The card names the source of record for each column.
-
2Week one · The exclusions
The first edge cases arrive
An out-of-office, and a reply from a colleague on the prospect's domain. Both go on the exclusions line the day they turn up, not later.
-
3Every Friday · The words
Ten minutes, nothing divided
Opt-outs, and the replies read as sentences. The counts go in the row. Nothing gets divided by anything this week.
-
4End of month · The read
Stage to stage, one segment
The five gaps get read in order, one segment at a time, against your own prior months. The card version goes in the row with the counts.
Every tool you add makes this worse before it makes it better. A new sequencer, a new CRM, a new booking link: each one arrives with its own version of every number. Decide the source of record for all six on the day the tool lands, not the month you notice they disagree.
Where the common advice is wrong
The pages that rank for this query do the same three things: name eight to twenty metrics, describe each in one sentence, and attach a benchmark range with no denominator behind it.
"Track these fifteen KPIs, and here is the benchmark range for each one."
- ✕More metrics is more visibility
- ✕A range tells whether your number is good
- ✕The metric name is the metric
- ✕Rates first, and counts never
"Keep six counts, write down how each one is counted, and compare next month to your own last month."
- ✓A list past six has lost the argument it is making
- ✓A borrowed range was counted with other operations
- ✓The name is the label; the formula is the metric
- ✓Counts survive a recount, and rates do not
What to realistically expect
Six counts and a card tell you where the motion leaks. They do not tell you whether the offer is right, and they do not produce volume.
A number that means the same thing in March as it did in January, a leak you can name instead of a month you can worry about, and an argument that ends because both of you counted the same way.
It will not make the list bigger, and it will not write a better offer. A measurement system measures. It does not produce companies worth writing to, and it never has.
What to take away
Fewer numbers, written down once, read on a schedule. Everything else on this page is detail underneath those five lines.
- 1 Six counts, not fifteen rates. A count survives being recounted.
- 2 Write the counting card before the first send, never after.
- 3 Read the formula under the label. They are not always the same fraction.
- 4 Every count needs the number that stops you gaming it.
- 5 Your baseline is four weeks of your own counts, not a benchmark.
Questions founders ask
What outbound sales metrics should I actually track?
What actually counts as a reply?
Do out-of-office replies count as replies?
What is a good reply rate for cold email?
Why do my sequencer and my CRM report different numbers?
What is the difference between a booked meeting and a held meeting?
How do I set a baseline with no benchmark to compare to?
Co-founder of Real Good GTM. He has been the first business hire and Chief of Staff at seed-stage B2B startups, building outbound pipeline before any playbook existed. He owns the systems and data side of the motion, which is where this guide comes from: the week somebody asks how it is going and two exports disagree by a third.
Connect on LinkedInOnce the counting is settled
These three pick up where this one stops: what a number can prove, which step-level numbers are worth keeping, and what a meeting costs to produce.
Outbound market learning
What a rate can and cannot prove once you have it, how wide its interval really is, and which denominators exist.
Read the guideOutbound sequences
Which step-level numbers are worth keeping, and how to tell a broken sequence from a broken list or a broken offer.
Read the guideCost per meeting
What a meeting actually costs to produce, including the operator hours nobody invoices, and which denominator decides the answer.
Read the postWant the outbound run, and the counting done properly?
Book a fit check. We'll look at your ICP, how you sell today, and what you can honestly measure at your volume. If outbound is not the right motion for your stage, we'll tell you that too.
Book a Fit CheckNo hard sell. No fake numbers. Real good work speaks for itself.