B2B lead sourcing: route decides the campaign
Lead sourcing is deciding which records exist for a campaign, and by what route you found them. The same company reached through a database filter, a job posting, and a funding feed arrives with three different things you are allowed to know about it, which makes it three different campaigns.
By Rahul Bageria, co-founder · Updated August 2026 · 23 min read
Written by operators who build these lists for seed-stage teams, not by a database selling seats.
What lead sourcing actually is
Sourcing decides who a campaign can possibly work on. Everything downstream, the copy, the timing, the offer, is capped by it.
B2B lead sourcing is finding the accounts and contacts a campaign will target, and recording where each record came from. It answers who exists, not what you know about them.
Sourcing is a route decision, not a lookup. The same account found by a database filter, by a job post, and by a funding feed hands you three different things you may open with.
Which means a sourced record is three fields, not one:
A human with a role in the decision, not a generic mailbox and not a title you inferred from headcount.
Which surface produced the row, and on what date. This is the only field that later tells you what is working.
Why this account, this week. If nobody on the team can write it, the row is filler and should not be there.
Four jobs people call one job
Sourcing, enrichment, ICP definition and market sizing get run as one blurred step. They fail in four different ways, which is why the cut matters the moment something breaks.
| The job | The question it answers | What it produces | How it fails |
|---|---|---|---|
| Sourcing | Who exists that we could talk to? | Accounts and named people, each carrying its origin. | A large, tidy list of the wrong companies. |
| Enrichment | What do we actually know about each one? | Filled fields: work email, mobile, headcount, stack. | Confident fields that quietly went out of date. |
| ICP definition | Who counts as a fit at all? | Criteria and red flags, written down before the query. | A wish list that describes nobody real. |
| Market sizing | How many of them exist in total? | A number for the board, and a ceiling on the motion. | A number nobody can act on this month. |
Sizing the market and building this quarter's list are different jobs, and only the second one has a send date. This guide assumes you already hold an ICP hypothesis. Without one, sourcing produces a large tidy list of the wrong companies, on time and under budget.
Where the discipline comes from
The account-first list is the idea that a target list is a designed artifact with stated criteria and stated red flags, not whatever a database query happened to return.
Define the key criteria of your ICP so you can focus your efforts on the most qualified prospects.
Trish Bertuzzi's Sales Development Playbook added the tiering in 2016: dream accounts worked by hand, bread and butter worked in a mix, a tier defined by a compelling event, and dead ends. That third tier is signal-based sourcing in a 2016 vocabulary. Everything invented since is plumbing.
The three routes to the same 500 records
Source the same 500 target accounts three ways and you get three campaigns, because each route hands you a different thing you are allowed to say.
| The route | How the record appears | What you legitimately know | What you can open with | Where it breaks |
|---|---|---|---|---|
| Filter-down | A database query on firmographics. | What they are. | A persona guess, dressed as relevance. | Everyone else ran the same query this morning. |
| Surface-up | They posted a job, listed an integration, spoke at an event, reviewed a rival. | What they are doing. | The thing they published, in their own words. | Slow, manual, and it does not scale by buying credits. |
| Signal-in | A round, a leadership hire, or a job change pushed them into your feed. | What just changed. | The change itself, and the deadline it created. | The window closes, and it closes for everyone at once. |
Which database you filter is a real decision, and the answer changes by geography, so it is worth reading how the major B2B databases actually differ on coverage before you commit a year of credits.
Before you write to a row, ask which route produced it. If the answer is "the database", your opener is a guess about a category. If it is "they posted this role in March", your opener is a fact they published themselves. Same account, different campaign, and the difference was decided at sourcing time.
The sourcing surfaces, ranked by what they tell you
Ranked by how much the surface tells you about the account, not by how easy the data is to buy. The database is one input, not the source.
The warmest rows you already own
Closed-lost deals past their cooling window, trials that went quiet, people who replied and never booked. Cheapest to source, warmest to open, and the surface founders skip most often because it feels like old news.
The pain, with a budget attached
A database tells you a company exists; an open job posting tells you the pain is funded, which is why it sits above the database in the surface order. The text was written by whoever owns the problem.
The stack, stated in public
App marketplaces, integration directories and partner pages say what a company already runs. That is the gap between guessing at a stack and reading it, and it decides whether you arrive as a replacement or an add-on.
The incumbent, named by the buyer
A public review names the tool a company uses and what annoys them about it. Slow to work through, and the only surface that hands you the competitor and the complaint in the same paragraph.
Budget, with a date on it
A round says money arrived and roughly what it is for. It is also the most crowded surface in outbound, because every alert fired the same morning, so treat it as a timing input rather than an angle.
The spine, not the source
A database gives you the firmographic skeleton: who exists, roughly how big, roughly where. It is the fastest way to get 500 rows and the weakest thing you will ever know about any of them. Last on this list on purpose.
Waterfall sourcing is not waterfall enrichment
The word "waterfall" is used everywhere in outbound and almost always means the enrichment kind. There is an equivalent discipline one step earlier, and it does not have a name yet.
Waterfall sourcing runs discovery surfaces in sequence to find records no single surface would have produced. Waterfall enrichment runs providers in sequence to fill a field on a record you already have.
Enrich your way out of a thin list
Run four providers across the same 200 rows until every email fills in.
- ✕The same 200 accounts, better formatted
- ✕Spend goes up, coverage of your market does not
- ✕You never find the account the query missed
Add a surface, then enrich once
Add job posts and marketplaces, reconcile, then enrich the combined set once.
- ✓New accounts no single query returned
- ✓Every row carries why it is there
- ✓Enrichment runs on a list worth filling
Everything after the record exists, completing and verifying the records you found, is a different discipline with different failure modes.
Build the exclusion list first
Before a single inclusion filter, write down who must never receive this campaign. It takes twenty minutes and it is the highest-value twenty minutes in the process.
-
1
Pull the rows you already own
Customers, open opportunities, and anything sitting in a live sequence. Export them by domain, not by contact.
Exact setting
stage in (Closed Won, Open Opportunity) OR sequence_status = active
-
2
Ask who is already in a thread
Investors, partners, advisers, and anyone a colleague is mid-conversation with. This part is not in the CRM.
Gotcha: One founder-to-founder thread you did not know about is worth more than the whole campaign it gets flattened by.
-
3
Add everyone who has said no
Past unsubscribes, hard bounces, and closed-lost accounts still inside their cooling window. Block at domain level, not address level.
Exact setting
unsubscribed = true OR hard_bounced = true -> block the whole domain
-
4
Freeze it as a file, not a memory
One suppression file, joined on domain and on address, checked again immediately before every send.
Gotcha: A re-source three weeks later will happily put every excluded row straight back, unless the file is joined again.
The worst outbound day I have had was a campaign that emailed a customer's champion, cold, asking whether they had a problem we were already solving for them. Twenty minutes of exclusion work would have caught it.
How many records is the right number
List size is a symptom, and it is diagnostic. Read the number as feedback on the definition upstream, never as a target to hit.
of business buyers are not in the market for a category at any one time.
John Dawes, Ehrenberg-Bass Institute, writing in 2021 for the LinkedIn B2B Institute. Dawes says plainly that the 95% figure is not meant to be a precise rule.
If nineteen rows in twenty cannot buy this quarter whatever you write, adding rows is the weakest lever you own. Forty thousand rows means you targeted a category. Ninety means you described one customer. Fix the definition, not the count.
One contact per account is under-sourcing
Sourcing one person per account is a decision to lose the account whenever that person leaves, goes quiet, or turns out not to own the problem.
people shape a complex B2B purchase, and each one arrives with their own research.
Gartner's 2019 buying-journey research puts 6 to 10 people in a complex buying group, each arriving with independently gathered information.
So source two to four per account, picked by role in the decision rather than by seniority:
The person whose week is worse because of the problem. Usually the easiest reply and the weakest signature.
The person the budget sits with. Slower to answer, and the only one who can turn interest into a decision.
Security, finance, or the owner of the system you plug into. Sourced now, or discovered in week six.
Two clocks: accounts decay slowly, people decay fast
A list does not go stale evenly. The account layer stays broadly true while the person layer rots underneath it, and one refresh cadence cannot cover both.
Refresh the two layers on different cadences and stop pretending one number covers it. Re-check the person layer before every send window. Re-check the account layer when the quarter turns, or when something on the account visibly moved. Anyone quoting you a single decay percentage is selling a refresh schedule, not describing your list.
Quarterly is honest
Headcount, funding stage, the stack, the problem itself. These do change, but slowly enough that a quarterly re-check is defensible. Re-sourcing the account layer every week is work that produces nothing.
Measured in a few years, not decades
Re-check the person layer before every send window: roles turn over inside a single campaign's shelf life. In the US, the Bureau of Labor Statistics, in Employee Tenure in 2024 released September 2024, put median tenure with a current employer at 3.9 years in January 2024, and 2.7 years for workers aged 25 to 34.
The shortest one on the row
A job post closes, a round stops being news, a review gets a reply from the vendor. The event that justified the record has its own expiry, and it is usually shorter than either of the other two.
The version of this that survives contact with a quarter is a list that refreshes itself on a schedule, with the person layer re-checked far more often than the account layer.
What "verified" actually means
A verified email is a claim about a moment, not a property of an address. Ask three questions of any verified flag: verified by whom, when, and against what.
On a catch-all domain, nothing can be checked. The server accepts anything, so "valid" there is an inference, not a test. Trust a provider's status labels over its accuracy headline: the labels are checkable against your own bounces, and the headline usually names no dataset and no period.
The inbox sets your list ceiling
Your sending capacity decides how many records you should source. Your credit balance has an opinion about it, and the opinion is wrong.
is the Postmaster Tools spam rate Google tells bulk senders never to reach, with 0.10% as the number to aim at.
Google's Gmail sender guidelines, whose bulk sender requirements have been in force since 1 February 2024.
Start from a daily cap you keep
Set a conservative daily send cap per mailbox and hold it. The exact number matters less than holding it, because a cap you respect is what keeps the spam rate flat while a new list is unproven.
Count only the warm ones
Count mailboxes that have finished warmup, on a domain that is not the one your customers email. A mailbox added last week is not capacity. It is a liability with an address attached.
A test window, not a year
Multiply by the sending days in the test, not in the quarter. You are sizing the list you will send to before you learn anything, and that is a far smaller number than people expect.
Steps eat the same budget
Every follow-up consumes the same daily capacity as a first touch. A four-step sequence roughly quarters how many new people you can start inside the same window. Whatever the arithmetic returns is your list.
Want a list built from the sources that fit your ICP, not one generic database?
Book a Fit CheckProvenance is a field, not a memory
Every record should carry where it came from and when. Without it you cannot judge freshness, cannot tell which surface produced replies, and cannot cleanly delete a source that turns out to be a problem.
When someone asks where you got their details, the answer is a lookup rather than an afternoon of archeology.
A provider or a surface that turns out to be bad can be pulled out by one filter, instead of poisoning the whole list.
"Which route produced the meeting" is answerable only if the route was recorded. That is the learning layer, in one column.
On every row, before the first send
5 checks
-
The surface is named, not the tool
Not "our data stack". The actual surface: this job board, this marketplace, this funding feed.
-
The date the row entered the list
So a two-week-old record and a two-quarter-old one stop looking identical in the same view.
-
The reason, written as one sentence
If nobody on the team can write it, the row is filler and it does not belong on the list.
-
Whether a human or an agent added it
An agent-written column is a claim. Marking it as one is what lets you audit it later.
-
The exclusion pass that cleared it
Stamp which suppression run the row survived, so a later re-source cannot quietly undo it.
The legal basis is set at sourcing time
The basis for a European campaign is chosen when the records are collected, not when you write the subject line. By then it is too late to satisfy it.
Consent is not the test in B2B
The CNIL, France's data protection authority, says B2B email prospecting can rest on legitimate interest rather than consent. Its guidance, last updated 10 June 2026, then attaches three conditions, and all three are set at collection.
All three, or none of it holds
The message must relate to the person's profession. They must have been informed, at collection, that their address would be used for prospecting. They must be able to object simply and free of charge, then and later.
You cannot retrofit condition two
The second condition is about something that happened before you ever saw the record. If you bought the list, you are taking a provider's word for it, and that word is only as good as where the provider got the data.
A provider's database can simply vanish
On 5 December 2024 the CNIL fined KASPR 240,000 euros over a database of about 160 million contacts built from LinkedIn. Its closure notice of 6 March 2026 records that KASPR deleted that database to comply.
Buy provenance, not volume
If your list depends on a provider whose coverage rests on a source it may not lawfully hold, you own a single point of failure that nobody priced. Ask where the records come from before you ask how many there are.
This is not legal advice. We are operators, not lawyers, we are based in Paris, and the rules differ by country and by how the record was collected. If a campaign turns on it, ask someone qualified.
LinkedIn: the best index you are least allowed to automate
LinkedIn is the best index of working people ever assembled, and its own terms rule out most of what sourcing advice tells you to do with it.
- ✓The only index people maintain about themselves
- ✓Role, tenure and team shape, in one place
- ✓Filters no general database can match
- ✓A message channel attached to the record
- !Scrapers, bots and crawlers are banned outright
- !Copying from aggregators and brokers is covered too
- !The commercial use limit is never disclosed
- !One restricted account can halt the whole process
LinkedIn's User Agreement, effective 3 November 2025, prohibits software, bots, crawlers and browser plugins used to scrape the service or download contacts. Its help pages describe a commercial use limit it will not disclose or lift, and cap one Sales Navigator search at 2,500 lead results across 100 pages, or 1,000 account results across 40 pages.
The consequence is architectural, not moral: build every search as slices small enough to finish inside those caps, and never build a sourcing process that stops working the day one account gets restricted.
What AI actually changed about sourcing
As of August 2026, agent columns changed what a filter can be. That is a real change at the discovery step, and it is not the change most people are selling.
A filter on facts no database stores: whether a pricing page shows a free tier, whether a careers page names a specific tool, whether a company sells into regulated buyers. A question, answered per account.
An agent column is a claim, not a field. It reads a page and writes an answer with the same confidence whether it was right or guessing, and nothing in the output tells you which.
Sample twenty rows by hand before an agent column is allowed to include or exclude anybody, and run that check again every time the prompt changes. Use agents to decide who is on the list, never what you say to them.
How we source a list, start to finish
This is the order we run it in, on every account. The order is the method: the tools underneath it change every year and the sequence has not.
-
1
Start from the exclusion file
Built before anything else, from your CRM plus a five-minute conversation about who is already in a thread.
-
2
Build the firmographic spine
One database pass for the accounts that could plausibly fit. The output is a starting set, never the list.
Exact setting
country + headcount band + funding stage + one hard qualifier you can verify
-
3
Layer the surfaces that carry timing
Funding, job changes, and hiring wired in, so the list carries timing, not just names. Built from the sources that fit your ICP, not one generic database that misses half your market.
Gotcha: Accounts found on two surfaces are not duplicates. Merge them and keep both reasons on the row.
-
4
Pick the people, then fill them in
Two to four per account by role in the decision. Then provider after provider until one returns a valid hit. Coverage no single source matches.
-
5
Hand it over as a living list
De-duped, fit-scored, synced to your CRM, and refreshed on a schedule so it never goes stale. A static export is old the day you buy it.
Exact setting
source_surface, sourced_on, reason, added_by, exclusion_pass on every row
We use Clay extensively and can help with Clay implementation, but the tool is not the method. The method is the order: exclusions, spine, surfaces, people, then a schedule that keeps the thing alive after we hand it over.
A worked example, start to send
Illustrative, not a client result. It carries arithmetic and method only: no reply rates, no meeting counts, no match rates. We report real numbers only when they are real.
Exclusions, then the spine
- Suppression file built and frozen
- One database pass on four criteria
- 4,000 accounts come back
- Nobody treats that as the list
The starting set, not the answer
Surfaces, then people
- Job posts and marketplaces layered on
- Accounts with a live reason kept
- Three people each, chosen by role
- Every row stamped with its surface
The list is now shorter and warmer
Cut it to what you can send
- Two warm mailboxes, 30 sends a day
- 15 sending days in the test window
- 900 sends, four steps each
- So 225 people get started
225 is the list. The rest is inventory.
Where lead lists go wrong
Almost none of these are data-quality problems. They are decisions made at sourcing time that only show up as a bad campaign six weeks later.
A customer, an investor, or a colleague's live thread gets a cold email. The cost lands on trust, not on the campaign.
The account dies the moment that person leaves or goes quiet, and you never learn whether the account was wrong.
Sourcing past your send capacity produces inventory, and inventory ages. The extra rows are stale before they are used.
A verification is a timestamp. Six weeks and a reorganisation later it is a claim about a company that has moved on.
A confident column that decides inclusion, never sampled by a human. Wrong rows look exactly like right ones.
Nobody can say where a row came from, so nobody can say which surface is working or remove a bad source cleanly.
Where the common advice is wrong
Almost everything written about lead sourcing is published by a company that sells data or sells lead generation. Four claims that do not survive a small team's reality.
"Buy the biggest database, filter by title, export everything your credits allow, and let volume do the work."
- ✕Start by buying the biggest database
- ✕More rows means more pipeline
- ✕Scrape LinkedIn to fill the gaps
- ✕Verify once, right before the campaign
"Write the exclusions, source three surfaces, and stop at what your inboxes can genuinely send to."
- ✓The database is the spine, not the source
- ✓Your inboxes set the ceiling, not your credits
- ✓LinkedIn's own terms rule that out
- ✓Verification is a timestamp, not a property
When to stop sourcing and start sending
Nobody who sells data will ever tell you to stop. Three conditions, and the first one that fires is the one that counts.
The arithmetic from the ceiling section returns a number. When the list covers it, you are finished, whatever the credit balance says.
If you would not personally email the last twenty rows you added, the useful part of the sourcing pass ended an hour ago.
Pick ten rows at random and say the reason out loud. If it does not come, the list has stopped being a list.
You learn more from a hundred sends than from a thousand more rows: replies correct the ICP, rows never do.
Key takeaways
- 1 The route you found a record by decides the campaign.
- 2 Write the exclusion list before the first inclusion filter.
- 3 Your inboxes set the ceiling on list size, not your credits.
- 4 Accounts decay slowly, people fast. Refresh them separately.
- 5 Every row carries its surface, its date, and its reason.
Questions founders ask
What is B2B lead sourcing?
What is the difference between lead sourcing and lead generation?
How many leads should be on a B2B outbound list?
Where can I find B2B leads without buying a database?
Is scraping LinkedIn legal?
Do I need consent to email B2B contacts in Europe?
How often should I refresh a lead list?
What does a verified email actually mean?
Co-founder of Real Good GTM. He has been a first business hire and Chief of Staff at seed-stage B2B startups, building the lists before he built the campaigns that ran on them. This guide is the sourcing half of how we work: what goes on a list, where each row came from, and when to stop adding to it.
Connect on LinkedInThe three questions this one hands off
How big the market is, whether to buy the list instead, and what to do with the list once it exists.
Total addressable market
Sizing the market is the job this guide hands off: how to count from the bottom up, and what the number is actually for.
Read the guideShould you buy lead lists?
The buy-or-build argument this guide deliberately does not make, with the deliverability and provenance costs priced in.
Read the postICP slice experiments
What to do once the list exists: split it into slices, send, and let the replies tell you which slice is the real ICP.
Read the playWant the list built for you, provenance and all?
Book a fit check. We'll look at your ICP, which surfaces your market actually shows up on, and whether outbound is the right motion for your stage. If it isn't, we'll tell you that too.
Book a Fit CheckNo hard sell. No fake numbers. Real good work speaks for itself.