What is B2B data enrichment?
B2B data enrichment completes and verifies records you already have: the missing work email, the mobile, the title, the headcount, plus a check that what is already in the row is still true. It does not find you a market, and it cannot rescue a list of the wrong companies.
By Kshitij Maheshwari, co-founder · Updated August 2026 · 23 min read
Written by operators who run enrichment for seed-stage teams every week, not by a vendor selling you a database.
What data enrichment actually is
Enrichment is a completeness-and-truth operation, not a discovery operation.
Data enrichment fills in the missing details on a contact or company, verified email, role, company size, tech stack, so a record is complete enough to act on, and confirms that what is already in the row still holds.
The fields split into three families, priced very differently:
Domain, headcount, industry, funding, tech stack. Mostly public, cheapest to fill, and the ones you should be scoring on before you spend anything else.
Title, seniority, function, tenure, who they report to. Mid-priced, and the first fields to go wrong the moment somebody changes job.
Work email, personal email, mobile. The expensive end of the table, and the only fields that can bounce, ring the wrong desk, or cost you a domain.
Enrichment has two jobs and they are not the same job: fill the blanks, then prove the blanks you filled are still true. A page that treats those as one thing is selling you a database.
What enrichment is not
Enrichment starts after the rows exist, so everything upstream of the rows belongs to a different job with a different failure mode.
- ✓The fields a row is missing
- ✓Whether what is already there still holds
- ✓Where each field came from, and when
- ✓How much bounce risk a send is carrying
- !Which companies belong in the rows at all
- !How big the market is, or what to count
- !Whether the person has any reason to care
- !Whether the tool you bought suits your market
Deciding which companies and people belong in the rows is the job that decides everything downstream, and it has its own method: building the list in the first place.
Coverage, accuracy and validity are three different things
Enrichment marketing uses these three words as synonyms. They are three separate measurements, and only one of them needs the real world.
The share that came back with something
Coverage is the share of your rows where a provider returned a value. A cascade raises coverage by construction, because it keeps asking until somebody answers. Nothing in that mechanism checks whether the answer is right.
Compliance with a rule
A verifier can prove an address is well formed, that the domain resolves, and that the mail server accepts it. Those are rule checks against a definition you control. This is what almost every tool in the category actually sells.
Correspondence with the real world
Accuracy needs a real value to compare against, not another database's opinion. Nothing outside the company can see whether that person still holds that job this morning, which is the thing you were actually asking.
How old the answer already is
Every field is a snapshot of the moment it was returned. Currency is how long ago that moment was, which is why a date beside a field is worth more than one more provider in the cascade.
Where the vendor claim collapses
A vendor can measure validity and deliverability. It cannot measure accuracy, because it has no access to the real value either. When a page says accuracy and shows you a validity number, that is not dishonesty. It is a different measurement.
Accuracy: the degree of closeness of data values to real values. Validity: the degree to which data values comply with rules.
That paper gathered 127 published definitions from nine authoritative sources and picked a preferred one for each, which is why it is worth quoting rather than paraphrasing.
How to read a "97% accuracy" claim
Every published accuracy percentage is a real number about something, and the only useful skill is working out what.
| The claim you read | What it usually measures | What you wanted to know |
|---|---|---|
| 97% email accuracy | Of the addresses returned and marked verified, this share passed the vendor's own validity check. | Of 200 companies you know, how many came back right. |
| 99% deliverable | Syntax, domain and mailbox checks passed at test time. | Whether the mailbox still belongs to that person next Tuesday. |
| 90% coverage | The share of a sample the vendor chose that returned any value at all. | The share of your ICP, in your geography, that answers. |
| Verified mobile | It passed that vendor's own process, which differs by vendor. | Whether the number reaches that person, not their old desk. |
| Real-time data | The lookup ran on request, not from a cached table. | How old the underlying record is. A fresh query can return a stale fact. |
Take 100 records where you know the truth: customers, past deals, people you have met. Run them through the provider and count the error-free rows. That number out of 100 is the provider's real score, the only accuracy figure about your list rather than someone else's sample. It costs an afternoon, and no vendor will suggest it.
Coverage is a property of your market, not of a brand, which is why this page runs no tool table. The shortlist question is answered separately in which B2B data tools are worth running.
Your data is worse than you think, and here is the measurement
Managers overestimate their own data, and somebody put a number on it with a method you can copy this week.
of newly created records carried at least one critical error.
Tadhg Nagle, Thomas Redman and David Sammon, Harvard Business Review, 11 September 2017, using the Friday Afternoon Measurement over two years.
The last 100 records you touched
Not a random draw across the database. The last hundred rows your process actually produced. No sampling theory, no tooling, no vendor involved in scoring their own work.
Ten to fifteen attributes that matter
For a contact list: name, company, domain, title, work email, mobile. Mark each record error-free or not. One critical error fails the whole record, which is the point.
Error-free records out of 100
That count is your score. The 2017 study ran the same instrument with 75 executives over two years, and nearly half of what they had just created was already broken.
The errors were in new data
Not aged data. Newly created data, already wrong about half the time. That moves the problem from decay to creation, and it is the half of the finding enrichment vendors quietly skip.
Why records go stale
Records break because people move, and you can size that from government data instead of the annual decay percentage this category likes to quote.
was the median time US wage and salary workers had spent with their current employer in January 2024.
US Bureau of Labor Statistics, Employee Tenure in 2024, released 26 September 2024. Workers aged 25 to 34: 2.7 years.
Every page in this category opens with an annual decay percentage. We chased the citation chain on the version everyone repeats and it dead-ends: a vendor blog citing a vendor blog citing a simulation page that credits research with no name, sample or method. Nobody can produce the study, so it is not on this page in any form.
Tenure carries the same point honestly. Sell into young companies with young teams and your person-level fields have a shorter half-life than any tidy annual rate suggests, and you can say so with a government release behind you.
You do not need a decay rate to set a refresh cadence. You need a date on the field and a rule about how old it can get before you stop trusting it. Ours is simple: a title older than six months gets re-checked before it appears in a first line, and an email older than a send gets re-verified.
The order of operations
Enrichment is the third of five moves, and a failure at move one is not recoverable at move three.
-
1
Source the rows
Decide which companies and which people belong on the list. Every step after this inherits that decision, good or bad.
Gotcha
A 90% match rate on the wrong companies is a faster, more expensive route to the wrong people.
-
2
Match to a real object
Resolve every row to one real company and one real person before you spend a single credit on it.
The test
One domain, one person, one row, and no second row hiding under a personal address.
-
3
Score, then enrich what survives
Fit-score on the cheap company fields you already have, cut hard, and only then buy person-level fields on the rows that made it.
-
4
Verify the slice, not the table
Run the verification pass on the segment about to be messaged, on the day it goes out.
-
5
Route it, with its provenance
Write the record into the CRM with the provider and the date attached to each enriched field, not flattened into one column.
Why the waterfall is the default now
Running one provider is a constraint choice in 2026, usually compliance or simplicity, rather than a starting point.
Cheapest confident source first
Providers are arranged so the cheapest source likely to be right gets asked first. The sequence is a cost decision, not a quality ranking, and it is yours to set.
First confident hit ends the run
A record stops at the first result the cascade is confident in. Everything below that level never runs on that row, which is what makes deep cascades affordable at all.
Misses are free, hits are not
Pay-on-success pricing means an empty lookup costs nothing. That removed the category's original failure mode and quietly installed a new one, which is what the cost section is about.
Coverage is not correctness
A cascade raises coverage by asking more sources. It checks no answer against the real world. Coverage went up. What accuracy did is a separate question nobody measured.
Clay's own published guide, updated 14 April 2026, states that in its 2025 work email benchmark no single provider cleared both 95 percent quality and 90 percent coverage. That is vendor documentation from a tool we use heavily, not an independent study, and the clearest statement of why cascades won: no source is good enough at both jobs.
That build is covered in how a waterfall is actually assembled, level by level.
Scope note: this section explains why the cascade won and where it stops paying. The credit math and the per-provider costs live in the waterfall piece, so nothing here restates them.
When to stop adding levels
The stop rule is arithmetic, not a provider count: keep a level only while its incremental hits are worth more than its cost per attempt.
Keep a level because coverage went up
You add a fourth provider, coverage moves from 71% to 74%, so the level stays on.
- ✕Coverage always goes up, that is what cascades do
- ✕Nobody priced the three extra points
- ✕The lift was once, the bill repeats monthly
Price the level before you keep it
Level four returns 3% more, at four times the cost per hit. Switch it off.
- ✓Measure incremental hits, never total coverage
- ✓Compare cost per hit to a contact's worth
- ✓Review monthly, switch levels off without ceremony
Where that curve flattens for a seed-stage team, and how many providers people actually end up keeping, is a numbers question with its own answer: how many data providers you actually need.
Mobile numbers are a different system
Phone data has its own sources, its own verification step and its own economics, and running it like email is the most expensive mistake in the category.
| Dimension | Work email | Mobile number |
|---|---|---|
| How it is found | Pattern plus a check against the domain. Close to solved in the US and Western Europe. | Assembled from fragmented sources. Coverage swings hard by country. |
| What verification proves | That a mailbox exists and accepts mail today. | That a number is live. Whether it reaches that person is a second question. |
| The second check | Not needed. The bounce is the check, and free. | Somebody calls the number and confirms the right person answers. Human cost, human speed. |
| Relative cost per hit | The cheapest lookup on every platform we run. | The most expensive by far, several times a work email. |
| When to run it | On the segment you are about to message. | Only on the segment you will actually call, after scoring. |
| What a mistake costs | A bounce, and a dent in your sender reputation. | A wrong number, a wasted dial, and a person who never agreed to a call. |
We do not cold call, so treat this as the cost shape rather than a calling motion. The rule holds either way: never run a mobile waterfall across a whole list.
Verify at send time, not at build time
A verification result is a timestamp, not a property.
A record you verified in March is not verified in August. The mailbox provider does not read your enrichment log; it reads the bounces from this send, on the day you make it. So the check that protects you runs on the segment just before it leaves. A quarterly pass over the table feels productive and touches none of that.
Two passes is the cheap version of this. One when the record enters the system, so you know what you are holding. One on the slice the morning it sends. The second pass is small enough to be almost free, and it is the one that catches the person who left in March.
Catch-all domains and what to do with them
Accept-all is not a gap in your verifier. It is the mail protocol working exactly as specified.
"Find a verifier that resolves catch-alls, then send with confidence."
- ✕Treat accept-all as a vendor quality gap
- ✕Buy a resolution pass and trust its label
- ✕Or delete the bucket to protect the bounce rate
- ✕Either way, no decision is ever written down
"Nothing outside can resolve it, so write a policy and measure what the policy returns."
- ✓Accept that no outside tool can see through it
- ✓Send the bucket separately, from a separate domain
- ✓Keep the slice small, read the bounce as data
- ✓Remember bigger, older companies cluster in here
A server may not fake a verification
RFC 5321, the SMTP specification published by the IETF in October 2008, says implementations "MUST NOT appear to have verified addresses that are not, in fact, verified", and that a server which disables verification "MUST return a 252 response".
Recipient checks can be deferred
The same document notes that where address validity "is deferred until after the DATA command is received, RCPT may return no information at all". An accept-all server is doing precisely that, on purpose.
Real and invented answer the same
So the domain returns an identical response for a real mailbox and one you made up. No verifier can resolve that from outside, and any product claiming otherwise is guessing with a confidence score bolted on.
It becomes a policy question
Which turns the whole thing from shopping into deciding: what is your rule for addresses nobody can resolve, and what will a small slice of them teach you this month?
Where AI enrichment helps and where it embarrasses you
AI columns are now standard for the fields no database holds, and unreliable for the fields they are asked to invent.
Reading a page that exists and pulling out a fact the page states. Does the careers page list a RevOps role. Does the pricing page show a free tier. What does the homepage say this company actually sells. The answer is checkable in one click, which is the whole test.
Anything the source page does not state. Inferred headcount, guessed revenue, a binary qualification call on a thin site. The model returns a confident answer either way, and confidence is not the same thing as a citation.
Ship a column nobody has read
Hi Sarah, saw you are scaling the RevOps team to twelve this quarter.
- ✕The source page never said that
- ✕One glance falsifies it
- ✕Worse than no personalization, because it is checkable
Sample 25 rows by hand first
Your careers page has two RevOps roles open in Berlin.
- ✓Quoted from a page you can open
- ✓Checked by hand before the column shipped
- ✓Never the only reason the email exists
What enrichment actually costs
Pay-on-success removed the old way to waste money on enrichment and quietly installed a better-hidden one.
Free, and no longer the problem
Nothing is spent when the cascade comes back empty. This is the cost teams still budget for and worry about, and it is the one that stopped existing.
The real bill, and it is invisible
Enriching ten thousand records to contact four hundred means most of what you bought is a row nobody will ever message. It never appears as a line item anywhere.
Every field is a running cost
A field costs a credit, adds a failure mode and gives somebody one more column to read. If it has not changed a message or a routing rule in a month, delete it.
Not all fields are priced alike
Person-level fields cost more than company-level ones, and mobiles cost the most by a distance on every platform we run. Running them across a whole list is the classic blowout.
The fix is an ordering, not a discount: score on the cheap fields you already have, cut to the segment you will genuinely contact, and only then buy the expensive ones.
Want the enrichment stack built and kept alive for you?
Book a Fit CheckEnrichment and the law
Enrichment is the moment one specific obligation is created, and it is the thing no page ranking for this term mentions.
Article 14 of the GDPR governs personal data "not obtained from the data subject". For enrichment that is not an edge case, it is the definition: a provider returned that work email, not the person it belongs to.
Enrichment is where the disclosure duty is created, and your first outbound email is when it comes due. The artifact that answers it is one column: which provider returned this field, and on what date.
Two duties follow. Article 14(2)(f) requires that you can say "from which source the personal data originate". Article 14(3)(b) sets the clock: where the data will be used to communicate with the person, the information is due "at the latest at the time of the first communication to that data subject".
That provenance column pays for itself twice more: it shows which level of your cascade is earning its place, and it lets you expire a field by age instead of re-running the whole table. Scope note: legal basis, and what each national regulator expects on top, is decided country by country and is not settled here.
Running it as a two-person team
The loop is small on purpose, and anything more elaborate than this at seed stage is a hobby with a subscription.
-
1
One pass when a segment lands
Run company-level fields once, on entry. Cheap, fast, and enough to score on. Nothing person-level yet.
-
2
Score, then cut hard
Fit-score on what you already hold and cut to the segment you will really contact this month.
The test
Would you email this person in the next four weeks? If not, do not enrich them.
-
3
Spend on the survivors only
Now run person-level fields, and only the fields a message or a routing rule genuinely consumes.
Gotcha
Mobiles only on the slice somebody will actually dial. Across a list they are the fastest way to burn a month of credits.
-
4
Verify the slice on send day
Re-verify the segment the morning it goes out, then write it into the CRM with the provider and date on every enriched field.
-
5
Half an hour, once a month
Check hit rate by provider, switch off the level that stopped earning, and delete the columns nothing consumed.
The monthly half hour is the step everyone skips, and it is the only one that compounds. Provider hit rates drift, a level that earned its place in March quietly stops earning, and nobody notices because the cascade still returns something. Look at the levels, never at the total.
Where enrichment programs die
Enrichment programs rarely die from bad data. They die from six habits, and every one of them is cheap to fix.
Credits go on rows nobody will message. The waste is invisible because a miss costs nothing and a hit still looks like progress.
Coverage tells you something came back. It says nothing at all about whether the something is right.
A quarterly pass over the whole table feels productive and protects nothing about the send going out on Thursday.
The most expensive field in the stack, bought for rows nobody will ever call. The fastest way to burn a budget.
A column that has not changed a message or a routing rule in a month is a credit, a failure mode and a distraction.
With no provider and no date on a field, you cannot answer a data-source question, price the cascade, or expire anything.
Where the common advice is wrong
Almost everything written about enrichment is published by a company that sells it, and five of the habits it teaches do not survive contact with a seed-stage budget.
"Buy the biggest database, clean it every quarter, and pick the verifier with the highest published accuracy."
- ✕Bigger database, better coverage
- ✕Clean the CRM on a quarterly cron
- ✕A published accuracy figure is a buying criterion
- ✕Skip every catch-all as a matter of course
- ✕Enrich everything, storage is cheap
"Score first, enrich the survivors, verify the slice on the day it sends, and benchmark the provider yourself."
- ✓Total size and coverage of your market are unrelated
- ✓The send-time check is what protects the domain
- ✓The denominator is undisclosed and differs per vendor
- ✓Skipping catch-alls is a default, not a decision
- ✓The cost is credits and attention now, not storage
What to realistically expect
Enrichment raises the ceiling on a good list and does nothing whatsoever for a bad one.
Fewer bounces, more rows you can actually message, and a shorter gap between deciding to contact somebody and being able to. Coverage on a market you have defined tightly is a solvable problem, and solving it is genuinely worth money.
Whether the company belongs on the list, whether the person has a reason to care this quarter, and whether your offer is worth a reply. No match rate has ever moved any of those three.
The record that just broke is the best reason to send. Somebody who changed jobs is not a maintenance ticket, they are the warmest trigger in outbound, and teams treating contact churn as cleanup delete their highest-intent signal every quarter. Enrichment tells you the field went stale. What you do next is a pipeline decision, not a hygiene one.
Key takeaways
If you keep five sentences from this page, keep these. Everything else on it is the reasoning behind them.
- 1 Enrichment completes and verifies rows you already have. It cannot choose the rows.
- 2 Coverage, validity and accuracy are three measurements. Only accuracy needs the real world.
- 3 Score first, enrich the survivors, verify the slice on the day you send.
- 4 Accept-all is a policy call, not a tooling gap.
- 5 Keep the provider and the date on every enriched field. One column, three payoffs.
Questions founders ask
What is B2B data enrichment?
What is the difference between data enrichment and lead list building?
How accurate is B2B contact data really?
Do I need waterfall enrichment or is one provider enough?
Should I email a catch-all address?
How often should I re-enrich my CRM?
Is enriching B2B contact data legal under GDPR?
Why do my enrichment credits run out so fast?
Co-founder of Real Good GTM. He has been the first business hire and Chief of Staff at seed-stage B2B startups, building outbound lists from an operator's chair rather than an agency script. The method on this page comes from running enrichment and verification on live campaigns every week, not from a spec sheet.
Connect on LinkedInFrom the model to the machinery
You have the model. These three cover the parts this page deliberately handed off: the list itself, the CRM it lands in, and the tools that run the cascade.
TAM Enrichment
A living target list rather than a one-time export: sourced, scored, enriched with a waterfall, and kept flowing into your CRM.
See the serviceCRM Hygiene
Where the enriched record ends up: keeping the CRM itself trustworthy, deduplicated and continuously enriched instead of cleaned once a quarter.
See the serviceBest waterfall enrichment tools
The shortlist this guide deliberately refuses to give you, compared honestly, including who should not buy one at all.
Compare the toolsWant this run for you, data and all?
Book a fit check. We'll look at your ICP, how much of it your current data actually covers, and whether outbound is the right motion for your stage. If it isn't, we'll tell you that too.
Book a Fit CheckNo hard sell. No fake numbers. Real good work speaks for itself.