Skip to content
GTM Guides

Cold email personalization

Personalization is not a setting you turn up. It is research effort you spend, and the only real question is which accounts earn it. This guide is the ladder of research depth, the rule that assigns it, and the arithmetic that decides how many accounts you can actually work.

By Rahul Bageria, co-founder · Updated August 2026 · 17 min read

The short answer Five lines, then the detail
What it is
Research effort spent on one account, so the message can say something that would not be true of a stranger.
The five rungs
Nothing, merged, segment relevance, account research, person research. A rung is named by what you had to find.
Who earns what
Account value times live signal, never job title. Title picks the reader inside an account you already chose.
The output nobody builds
A research pass can legitimately end in do not write. That decision is a deliverable, not a failure.
The catch
Automate collection. Spot-check extraction. Never automate the call on whether a fact obliges anybody.

Written by operators who run this motion for seed-stage teams, not by a vendor selling a research agent.


What personalization actually costs

Every ranking page treats personalization as a dial. It is a budget. Effort you spend on one account is effort no other account gets that week.

Definition

Personalization is research effort spent on one account, spent so the message can say something true of them that a stranger would not recognize.

Also called personalization at scale · the glossary entry

Three things get spent, and only one of them is yours:

Your hours

The input you control. Hours on research are hours not spent writing, sending, or answering the replies.

Their attention

The input you are bidding for. Herbert Simon, in 1971: a wealth of information creates a poverty of attention.

The proof of effort

A fact anyone's agent could fetch in a second proves nothing. Cheap effort cannot signal effort.

Personalization is not something you do. It is something you spend, and the only real question is who earns it.


Relevance is not personalization

Relevance is about the reader's situation. Personalization is about the reader. One can be generated, the other has to be found. Our cold email copywriting guide sets that split out in full.

The through line

This page starts one step earlier. Before anything can be relevant, somebody has to go and find the thing that changed, and somebody has to decide which accounts are worth the looking.


The ladder

Five rungs of research effort

A rung is named by what you had to find. Each one costs more than the last and entitles you to say more than the last.

Rung What you had to find Cost per account What it lets you claim Can the reader tell?
0. Nothing Nothing. None. Nothing. Instantly.
1. Merged A column you already had. None. It is a column. Still nothing. Yes, and it is now the default assumption.
2. Segment relevance One reason, verified true of twenty to sixty accounts. An afternoon per segment, then pennies. What this group is on the hook for. Only if the reason is wrong.
3. Account research One fact that changed what this company owes itself. Five to ten minutes. What this account now has to deal with. Yes, and this is the rung that reads as attention.
4. Person research Why this person, and not their colleague, owns the thing that changed. Ten to twenty minutes, and rising. Why you are writing to them specifically. Yes, and it is the easiest rung to fake.

The rungs and the minute bands are how we run the motion, not a measured benchmark. No published dataset measures cold email reply rate by depth of research, so no rung here carries a rate.


Why a rung is not a merge field

Counting merge fields is the wrong axis. The right one is where the fact came from, and what it would have cost anybody else to get it.

The right axis, in four beats
1
The habit

Levels get counted in placeholders

Every ranking page defines its levels by how many personalized slots the message carries. Add a name, add a company, add an image, call the result hyper-personalized.

2
The evidence

More placeholders sold worse, not better

The one randomized test of that idea, run on an Italian retailer's triggered email in 2023, found the heaviest level converted worst. Read the direction, not the size.

3
The test

Name the rung by where the fact came from

Ask two things of every line. Where did this come from, and what would it have cost anybody else to get it? Six merged fields off a list you bought is still rung 1.

4
The move

Merge the fields when the argument is the segment's

If the sentence is true of all sixty accounts, a merge field is the honest way to send it. Spend the researched facts on the ten where something actually happened.

The rungs are not cumulative

Rung 4 does not contain rung 3. A fact about a person, with no fact about their situation underneath it, is trivia at a higher price.


What each rung actually buys

Three of the rungs have something measured behind them. Read what each one buys before you decide which rung an account is worth.

Rung by rung
Rung 1

Merged buys a glance, not a reason

A name in the subject line does lift opens a little, measured in 2018 on opt-in consumer email. It buys the open and nothing after it, so treat rung 1 as postage rather than personalization.

Rung 2

The best-measured rung there is

Hunter's State of Email Outreach 2026, on 31 million emails its users sent in 2025, puts sequences of 21 to 50 recipients at a 6.2% reply rate against 2.4% past 500. Vendor data, and smaller lists tend to be better-aimed lists as well.

Rungs 3 and 4

Company beat person with senior buyers

Gong Labs, updated March 2026, puts company-level personalization ahead of individual-level with directors and above. Write about what the company now has to deal with before you write about the person.

Source: Gong Labs, updated March 2026, vendor data, so the ordering travels and the multipliers do not. Nobody has published the reply gap between rung 3 and rung 4 on cold email.


The ladder test

A rung counts only if it changed a sentence you could not have written without it. Delete the fact, re-read the draft, and see whether it survives.

Don't

Climbed, and nothing changed

Saw you are hiring a RevOps lead. We help teams like yours clean up their CRM.

  • Delete the hiring fact and it still reads
  • The claim was never about them
  • Ten minutes spent, nothing bought
Do

Climbed, and the sentence needed it

Your RevOps posting lists both Salesforce and HubSpot, so the first job is picking one.

  • Delete the fact and the sentence collapses
  • Names what they now have to decide
  • True of a handful of accounts, not sixty

Constructed examples written for this guide, not messages from a campaign.


Account tiering: who earns which rung

Tier the accounts, not the personalization. Two variables multiplied: what a deal is worth, and whether anything has actually happened. Title is not one of them.

Tier A · rung 3 or 4
  • Worth a real deal, and something changed
  • Five to twenty minutes of research each
  • Ten to thirty live at once, as a queue
Tier B · rung 2
  • Worth writing to, but nothing has happened
  • Shares one verified reason with twenty to sixty others
  • Where most of the work and most of the return sit
Tier C · rung 0
  • !Neither the value nor the event is there
  • !The output is a decision to skip
  • !A cheaper email is not the cheaper option

Tiering effort by account worth is not a cold email idea. Bev Burgess and Dave Munn's A Practitioner's Guide to Account-Based Marketing opens its account-selection chapter with the line that not all accounts are equal.

Operator note
Learned the hard way

Decide the rung once a week, per tier, and never at the moment of writing. Depth chosen while drafting is depth chosen by whatever energy you had that morning, and that is the one moment it is reliably chosen badly.

RB
Rahul Bageria
Co-founder, Real Good GTM

Tier C's output is not a cheaper email

Every founder builds tiers A and B, then quietly emails C anyway, because the list already exists.

Research has two legitimate outputs: a reason to write, and a decision not to. Most teams have only ever built the first one, which is why they always find something, and why what they find is usually trivia.

A pass that disqualifies nothing means the bar is on the floor.


The weekly research budget

You do not decide how personalized your emails are. You decide how many hours a week go to research, and the tiering decides who gets them.

112
touches in a median rep's day

is the load already on the calendar, before a minute of it goes on research.

As of August 2026

The Bridge Group, Sales Development Models, Metrics and Compensation, February 2025, on mostly North American companies.

The budget, in three beats
1
The load

Start from four minutes a touch

An eight-hour day is 480 minutes. Divide it by those 112 touches and each one gets about four, before research, list building, CRM logging or lunch.

2
The budget

Budget an afternoon per segment, ten minutes per account

Rung 2 costs one afternoon a segment and then pennies each. Rung 3 costs five to ten minutes an account, rung 4 ten to twenty. Multiply by your Tier A queue for the week's number.

3
The read

Ten minutes a prospect is a choice

The standard advice of ten to fifteen minutes per prospect is not a tempo. It is a decision to work far fewer accounts, and nobody recommending it says so.

Those are teams with dedicated reps, not a two-person founder-led one, so four minutes a touch is a ceiling rather than a picture of your day.

Want the tiering built against your own account list?

Book a Fit Check

What you are looking for, and where to look

One test decides whether a fact is worth anything. Does it create an obligation, a deadline, or a decision for them?

Definition

A qualifying fact changes what the account is on the hook for. "They posted about RevOps" creates nothing. "The RevOps posting puts CRM consolidation in the first ninety days" creates all three.

  1. 1

    Their own published statements

    Job postings, careers page, changelog, docs, pricing, status page, filings. Cheapest to check, hardest to argue with, and the source almost nobody opens first.

  2. 2

    The event record

    Funding, acquisitions, executive hires, launches, migrations, regulatory dates. Hiring is the one most founders can act on without buying anything, and hiring signals covers how to read a posting.

  3. 3

    Their own public professional statements

    This is rung 4, and it earns its minutes only when it explains why this person, and not their colleague, owns the thing that changed.

  4. 4

    Inference. Stop before this one.

    Inference is where research turns into invention. A public signal is fair game and private data is not, and personalization without being creepy owns that line.


When to stop, and what disqualification teaches

Set a timer per tier and honor it. When it runs out and nothing has changed for them, the answer is not to look harder.

What the skipped accounts tell you
  • !A pass that never disqualifies is decoration with a stopwatch
  • !"Already solved it" is a positioning problem
  • !"Cannot find anything" is a segment problem
  • !"Found something, but it obliges nobody" is an offer problem

Log the reason every time an account gets skipped. Read the column monthly and the segment definition improves without a single reply arriving.

Operator note
The part people skip

The timer is the part everybody drops first. A stopwatch running on one account is harder than it sounds, and it is the single change that most improves the output, because it forces you to notice how often nothing is there.

KM
Kshitij Maheshwari
Co-founder, Real Good GTM

What to automate, and what never to

Three jobs hide inside the word research, and they have three different answers.

The job Automate it? Why
Collection: finding and fetching the page Completely. A machine reads sixty careers pages faster and far more patiently than you do.
Extraction: turning a page into a fact With a spot-check. Right most of the time is not the same as right, and the misses are quiet.
Judgment: whether the fact obliges anybody Never. This is the whole job, and it is the part being sold as solved.

Clay's own documentation describes Claygents as agents "designed for judgment-based GTM work like account research, lead scoring, outbound copywriting, and persona classification". Checked August 2026, it carries no accuracy caveat and no requirement to review an output before it reaches a prospect.

!
Caution

Detected AI writing costs sincerity, not polish

In a 2025 study led by Peter Cardon at USC, 1,100 professionals rated manager emails with the machine-written parts marked. Sincerity fell from 83% under light assistance to between 40% and 52% under heavy. Professionalism barely moved.

Do this instead
Let a machine collect and extract. Type the sentence that makes the claim yourself.

That study covers internal manager-to-employee email, not cold outbound, and participants were told which parts were machine-written, so it sizes the penalty once AI use is believed.


How we would run it

A worked example: one week, three tiers

An illustrative walkthrough of the method, not a specific client result. We report real numbers only when they are real.

  1. 1
    Monday · The budget

    Four hours, and that is all

    Four hours go on the calendar for research this week. Everything below is spent inside that block, and nothing borrows from next week.

  2. 2
    Monday · The sort

    Sixty accounts into three tiers

    Eight have a live event and real value. Thirty-four share one verified reason. Eighteen have neither, and they are the ones nobody wants to sort.

  3. 3
    Tuesday, Wednesday · Tier A

    Ten minutes each, timer running

    Two of the eight get disqualified on the timer: one had already solved it, one had nobody who owns it. Both reasons go in the log.

  4. 4
    Thursday · Tiers B and C

    One argument, and eighteen skips

    One argument gets written once for the thirty-four. The eighteen get a skip and a logged reason. Six accounts end the week with sentences only they could receive.


Where personalization programs die

Programs rarely die from bad data. They die from four decisions, and every one of them feels like effort at the time.

Climbing the wrong rung

Person facts feel like work, so teams reach for rung 4 while rung 3 sits there cheaper, more durable, and better with senior buyers.

Burning a rung-3 sentence

If the fact is true of all sixty accounts, it belongs in the segment argument. Spend the expensive minutes on something that is not.

A pass that never disqualifies

Researchers who always find something are not researching. The empty disqualification column is the proof, and it is easy to check.

Deciding depth per email

Depth chosen at the keyboard drifts to whatever the morning allowed. Depth belongs to the tier, decided once, away from the draft.


Pushback

Where the common advice is wrong

Most personalization advice is written by companies selling personalization, and the numbers behind it do not survive being opened.

The common advice

"Hyper-personalize every email, spend ten to fifteen minutes a prospect, and watch the replies come."

  • More placeholders means more personalization
  • Ten to fifteen minutes on every prospect
  • Point an agent at it and personalize at scale
  • Person facts beat company facts
What actually works

"Fix the hours you can spend in a week, then decide which accounts earn them."

  • The only randomized test found the highest level worst
  • Ten minutes each is a decision to work fewer accounts
  • Agents collect and extract; judgment stays human
  • Company-level beat person-level with directors and above

What works now

What changed, and what did not

Current as of August 2026, and dated on purpose, because all of it moves. The ladder is craft. This section is weather.

Since 2023 What moved What it does not change
Everyone has the same agent A fact any agent could fetch stopped proving effort, because provenance was part of what the line signaled. A fact about a person is still not a reason to write.
The AI penalty got measured Peter Cardon's 2025 USC study put a number on what happens once a reader believes the writing was mostly machine. Light assistance carried almost no penalty. Heavy composition did.
The direction moved to the company Gong's dataset puts company-level personalization ahead of person-level with directors and above. The delete test predates all the tooling and is unaffected by it.
What did not change

The obligation test and the delete test were true before any of this shipped, and they are true after it. Anybody selling a 2026 framework for either one is charging you for arithmetic you already own.

Operator note
What we changed

The thing we changed this year is where the agent sits. It fetches and it extracts. The sentence saying what the fact obliges them to do is still typed by a person, because that sentence is the only part the reader is actually reading.

RB
Rahul Bageria
Co-founder, Real Good GTM

What to realistically expect

The ladder changes who you write to and what you are entitled to say. It does not change whether the offer is worth answering.

What it moves

Fewer accounts, better aimed, carrying sentences that could not have been written about anybody else. On current platform data the tightest sequences reply best, which is the segment rung turning up in somebody else's numbers.

What it does not move

A rung-4 email against a weak offer earns a polite, specific no from somebody who now knows exactly why they do not need you. Faster rejection, better informed. The outbound offer is the layer underneath.

The through line

Your research budget is not a gift to the reader. It is a bid for something they are already rationing, and the bid is read against every other message arriving that morning.


Key takeaways

Five lines that close the argument, in the order a Monday actually needs them.

Key takeaways
5 points
  • 1 Personalization is a budget you spend, not a setting you raise.
  • 2 A rung is named by what you had to find, never by fields merged.
  • 3 Tier by value times live signal. Title is a filter, not a tier.
  • 4 Research has two outputs, and one of them is a decision to skip.
  • 5 Automate collection, spot-check extraction, never automate judgment.

FAQ

Questions founders ask

What is cold email personalization?
It is research effort spent on one account, spent so the message can say something true of them that would not be true of a stranger. It is a budget rather than a setting. The useful question is not whether to personalize, but which accounts earn how much of your week.
How much time should I spend researching each prospect?
Not a per-prospect number, a per-week number. The Bridge Group's 2025 report puts median sales development activity at 112 touches a day, which leaves about four minutes a touch in an eight-hour day before any research at all. Decide the weekly research hours first, then let tiering decide who gets them.
Is more personalization always better?
No, and the only randomized test of the question says so. Run on an Italian retailer's triggered email in 2023, it found the heaviest placeholder level converted worst on purchase. That was consumer retail, so read the direction rather than the size. What you are counting is placeholders, and placeholders are not what earns a reply.
How do I decide which accounts get deep research?
Account value times live signal, never job title. Accounts where a deal is worth enough and something has changed get real research. Accounts that share a verified reason but have no live event get segment work. Accounts with neither get nothing, which is a decision rather than a failure.
What should I be looking for when I research an account?
A fact that changes what they are on the hook for. If it creates no obligation, no deadline and no decision for them, it is trivia however hard it was to find. "They posted about RevOps" fails that test. "The RevOps posting puts CRM consolidation in the first ninety days" passes it.
Can I use AI to personalize cold emails at scale?
Automate collection, spot-check extraction, never automate judgment. In a 2025 study led by Peter Cardon at USC, 1,100 professionals rated internal manager emails with the machine-written parts marked, and sincerity fell from 83% under light assistance to between 40% and 52% under heavy. Let a machine fetch and extract, then type the sentence that makes the claim yourself.
When should I not send the email at all?
When the timer ran out and nothing you found changes what they are on the hook for. Sending anyway produces a rung-1 email to an account you already judged cold. Log the reason for the skip instead, because those reasons are the only place a positioning problem, a segment problem and an offer problem look different.
Rahul Bageria, co-founder of Real Good GTM
About the author
Rahul Bageria

Co-founder of Real Good GTM. He has been the first business hire and Chief of Staff at seed-stage B2B startups, building outbound pipeline before any playbook existed. This guide is the budget he had to invent with sixty accounts, two people, and no afternoon left to give any of them.

Connect on LinkedIn

Keep going

From the research to the message

You have the budget and the tiers. These three cover what you do with the fact once you have found it.

Want the tiering built for your ICP?

Book a fit check. We'll look at your account list, which events actually fire in your market, and how much research your week can honestly carry. If outbound is not the right motion for you yet, we'll say so.

Book a Fit Check

No hard sell. No fake numbers. Real good work speaks for itself.