Automate prospect research without slop
You can automate the finding and the fetching. You cannot automate deciding what a fact means for your offer. The way to keep that honest is to make every automated answer arrive with the line it came from, its link and its date. Here is the test, the shape, and the check.
By Kshitij Maheshwari, co-founder · Updated August 2026 · 11 min read
Spot the four ways a research line goes wrong
Almost every bad automated line is one of four shapes, and all four land with the one person alive who can falsify them instantly. Learn the four and you can catch one without reading two hundred emails.
| The shape | How it reads in an email | The check that catches it | What the check costs |
|---|---|---|---|
| The confident invention | "Saw you opened a Berlin office last month." | Open the link the step gave you and search that page for five words of the quote. | Opening one page. |
| The horoscope | "Noticed you are scaling the sales team this year." | Put a competitor's name in and read it again. Still true means it says nothing. | One swap, before you write. |
| The true and pointless | "I see you run a help center on Zendesk." | Ask what the fact makes them responsible for this quarter. Nothing means it is trivia. | One question, in your head. |
| The wrong company | "Following your move to a four-day week." | Match the domain on the source link against the domain you are sending to. | One look at two addresses. |
No published measurement exists of how often an automated research line is wrong: no vendor publishes an error rate for its own research step, and no research institution has measured one. The four shapes come from the mechanism, and from the check that catches each.
Ask whether someone could point at it
One question decides what you automate, and you can run it on the step in front of you before you finish reading this paragraph.
The show me the line test: could a person answer this by pointing at one line on one page? If yes, a machine can. If the answer only exists after somebody decided what that line means for your business, that part stays yours.
- •What did their changelog say shipped last month?
- •Which payment providers are named on their integrations page?
- •How many locations does their contact page list?
- •Which customer does their newest case study name?
- •Is that release a problem for them or a victory lap?
- •Does the thing they added replace what you sell, or feed it?
- •Is the person who announced it the person who owns the budget?
- •Does any of it earn you a first line, or is it merely true?
How much research an account is worth is a different question, and our cold email personalization guide settles it. This page starts after that decision, and asks what the automated step has to hand back.
If the step cannot show you the line, you have not automated research. You have automated writing.
Make the answer arrive as a row, not a paragraph
Same account, same question, two output shapes. One of them can be checked across two hundred rows and the other cannot.
The question asked of both: did this company ship anything in the last sixty days that touches what we sell?
Ask for a summary
Acme has been investing heavily in automation and recently expanded its platform, positioning them well for growth.
- ✕Nothing in it can be opened
- ✕Checking it means doing the research again
- ✕It will read like this on every row
Ask for four fields
Yes / "Scheduling now runs on our own queue" / acme.com/changelog / 14 July 2026
- ✓The claim and its receipt travel together
- ✓Checking is opening one link
- ✓An empty answer stays visibly empty
An illustrative walkthrough of the method, not a specific client result. We report real numbers only when they are real.
A changelog is the cheapest surface to test this on, because it is dated, public and theirs. What a release actually entitles you to say is in our product launches signal.
Tell it what to do when it finds nothing
One sentence added to the instruction, and your empty rows stay visibly empty instead of quietly becoming fiction.
A step that is never allowed to say "I found nothing" will always find something. Columbia's Tow Center tested AI search tools in March 2025 and found them generally bad at declining to answer questions they could not answer accurately. They give an incorrect or speculative answer instead.
Write the permission into the instruction: if the page does not state it, answer nothing found and leave the quote blank. Different tools, same habit. Staying quiet is not the default, so you have to ask for it.
Check the lines, not the emails
A column of short quoted lines can be read in one sitting. A stack of finished emails cannot, and the wrongness lives in the lines anyway.
-
1
Open the rows that said yes
A random sample is mostly rows that found nothing. The expensive errors sit in the confident positives going out this week.
Exact setting
Filter: answer is not empty. Sort: send date, this week first.
-
2
Read the quote against the link
You are checking one thing: whether the page says the line. Whether a bought field is right at all is a different job, and our data enrichment guide owns it.
-
3
Hand it to the other person
Whoever wrote the instruction reads the output for what they meant. Somebody who never saw it reads it for what it says. On a two-person team, swap.
There is no budget argument for skipping this. Clay's own product documentation, checked August 2026, says you can test up to 10 inputs at a time for free and that test runs do not cost credits.
Want the research layer built so you can check it in a sitting?
Book a Fit CheckOpen the link before you believe the quote
A link is not a check. Clay's documentation says its research agents decide which sources to consult, and that decision never reaches you unless the output carries the link it picked.
"It came back with a source attached, so somebody has checked it."
- ✕A link is a claim about a page, not proof of one
- ✕You never saw which page it decided to believe
- ✕The address can point at a page that never existed
"It is a place to look, and looking costs one page open."
- ✓The quoted line is either on that page or it is not
- ✓Five words and a find-on-page settles it
- ✓A row you cannot settle is a row you delete
Somebody graded the answers
The BBC and the European Broadcasting Union graded how AI assistants answer news questions and published the result in October 2025. Of the answers, 45% carried something that could materially mislead the user.
The source that does not say it
One failure they name is an answer that cites a source which does not contain or support the claim. The link works, the page is real, and the sentence is not on it.
The link that never existed
They also name sources where the page or the exact address never existed at all. Nothing in the output tells you which kind you are holding, which is why the check is opening the page.
None of that was about sales email. It is about what a citation is worth when nobody opens it, and your research columns run on exactly the same footing.
The dangerous sentence is not the clumsy one. It is the fluent one about a company that is not theirs.
Decide where a wrong answer is allowed to land
Truth is one of two questions a line has to pass before it sends. The other is permission, and personalization without being creepy owns that one.
The same wrong answer costs you differently depending on where you put it. In a filter it costs one skipped account, quietly. In a sentence you sent, it costs the account and your credibility with the one person who can falsify it.
So the column you trust least belongs in the list logic, not in the first line. Same imperfect answer, and the worst case becomes a skipped account instead of a burned one.
Count what checking costs before you count what the tool costs
The subscription is the visible number and the small one. What you are buying is a pile of answers somebody now has to look at, and at two people that somebody is you.
Ask how many rows a week you can honestly open. That number is your sending volume, whatever your plan says.
A quoted line and a link can be scanned. A paragraph has to be researched again, which is why nobody does it.
If you cannot check what you are sending, send fewer emails to accounts you can say something true about.
Automation swaps a limit on how many accounts you can research for a limit on how many you can check.
The three ceilings a seed team runs into are in our seed-stage outbound playbook. Research is the one automation moves rather than removes, and that move is the whole reason the check has to exist.
We size a campaign by what we can read, not by what we can enrich. When the two disagree, the list gets shorter. It is the least popular rule we run and it has never once been the wrong call.
Four ways teams break their own check
Every one of these is a team that built a check, ran it once, and then quietly stopped.
Every edit makes it a new step, and last week's check does not carry over. Clay's documentation says the same thing mechanically: editing a saved instruction does not change the columns already running on the old one.
Reading the emails that went out is not a check, it is a post mortem. The line has already landed with the person who can falsify it.
You read the output for what you meant it to say. That is not a character flaw, it is what writing the instruction does to you, and swapping is the entire fix.
Twenty random rows are mostly blanks and easy yeses. The expensive mistakes live in the confident positives, and a random sample barely touches them.
Getting their own company wrong
A line that is wrong about the reader's own business is the mistake outbound does not recover from. They know it is wrong in the time it takes to read it, and the more specific you were, the more certain they are that nobody checked.
- 1 Automate a step only if a person could answer it by pointing at a line.
- 2 Ask for four things back: the answer, the line, the link, the date.
- 3 Make "nothing found" a permitted answer, in writing, or you get a sentence.
- 4 Check lines rather than emails, and check the confident yeses first.
Questions founders ask
Can I automate prospect research without the emails sounding like a robot wrote them?
How do I know if the AI made up a fact about a prospect?
What should an automated research step actually give me back?
How many rows should I check by hand?
Why do my personalized lines all sound the same even though they are true?
My tool cites a source. Is that enough?
Co-founder of Real Good GTM. He has been the first business hire and Chief of Staff at seed-stage B2B startups, building outbound pipeline before any playbook existed. He writes and checks these research columns every week, and this post is the shape he makes them hand back before anything they wrote is allowed to send.
Connect on LinkedInThe questions on either side of this one
How much research to do before this page starts, what the data layer underneath owes you, and the other check a line has to pass.
Cold email personalization
How much research an account is worth, what each level of effort buys you, and when to stop.
Read the guideData enrichment
The layer underneath: what coverage and accuracy claims really mean, and what a field costs.
Read the guidePersonalization without being creepy
The other pre-send question. Not whether the fact is true, but whether you may say it out loud.
Read the postWant a research layer you can actually check?
Book a fit check. We'll look at what your columns are being asked to answer, what they hand back, and whether the volume you are planning is a volume anyone could check.
Book a Fit CheckNo hard sell. No fake numbers. Real good work speaks for itself.