← BUILD LOGSCORING EVIDENCE
BUILD LOG · THE WIN/LOSS BACKTEST

I backtested our lead score against real won and lost deals. The number-one predictor of a win was not even in the model.

Everyone trusted the fit score. So I tested it against the deals we actually won and lost, and the signals that really separated them were behavioral, not firmographic, and the strongest one was not weighted at all.

2.67x
win-lift from the strongest behavioral signal
3.56x
that signal stacked with an ICP threshold
16 to 0
won-lost record of every account that reached Hot or Strike
1.13x
the fit score everyone trusted, barely discriminating
THE STORY, FROM THE SEAT

Here is one I actually shipped. What happened, what I did, and the stack if you want to run it yourself.

Every scoring model is a set of assumptions until you test it against reality. Ours assumed fit predicted wins.

But when I pulled the won and lost deals and measured what actually separated them, the fit score barely moved the needle, and several accounts we lost had scored in the 90s on fit. A model that scores your losses as highly as your wins is not scoring. It is decorating.

The Sales Operator
Keep Building,
Heath
FOUNDER, THE SALES OPERATOR
01 / HOW WE APPROACHED THE PROBLEM

Ran the loop.

Every scoring model is a set of assumptions until you test it against reality, and ours assumed fit predicted wins. The test was to pull the deals we actually won and lost and measure how often each signal showed up in each. Lift, not vibes. The number-one predictor of a win was not even in the model.

  1. 1
    Measure lift, not opinion

    For every signal, I compared how often it showed up in won deals versus lost ones. Lift, not vibes. A signal that appears equally in wins and losses is noise, no matter how good it feels.

  2. 2
    Find the predictors that were not in the model

    The strongest discriminators were behavioral, how the buyer actually worked, not the firmographics we had been scoring on. The single best one carried a 2.67x lift and was not weighted at all.

  3. 3
    Stack the signals that compound

    One behavioral signal plus an ICP threshold stacked to a 3.56x lift. The combination beat either alone, which is the whole argument for keeping signals separate so they can compound.

  4. 4
    Retire the flatterers

    The fit score and CRM presence both hovered near 1.1x, barely better than a coin flip. I stopped letting them carry weight they had not earned.

02 / WHAT IT TOOK CROSS-FUNCTIONALLY

Nobody ships this alone. Here is who had to move.

The backtest lived or died on data other teams owned: the behavioral events and the deal history had to be open before lift could be measured honestly.

PRODUCT AND DATA
The behavioral event access.

The strongest discriminators were behavioral, how the buyer worked inside the product, and that signal lived in the product warehouse. The backtest only ran because the events were opened up to the revenue side.

REVOPS
The deal history the model was tested against.

Wins and losses had to be clean and complete for lift to mean anything. A backtest against messy outcomes just decorates a different way.

03 / HOW WE TURNED IT INTO A SALES MOTION · THE PEOPLE PART

The backtest found the real predictor. Making the debrief a team habit is what turned it into wins.

A model that knows why deals close is worth nothing if the team never revisits a closed deal. The leadership half was turning win-loss from a spreadsheet into a standing ritual the whole team learned from.

THE FRAMEWORK BEHIND IT
The win-loss debrief, run on PLAN

Every closed deal, won or lost, gets read back against the five PLAN questions: was the problem real, were priorities lined up, were decision dynamics mapped, was there a compelling event, were next steps mutual. The pattern across debriefs becomes the coaching.

  1. 01
    Ran a debrief on every closed deal, not just the losses

    Wins teach as much as losses when you ask why honestly, and separate the great demo from the warm intro.

  2. 02
    Turned the top predictor into a qualification question

    The signal the backtest found became something reps checked for on live deals, not just a finding in a deck.

  3. 03
    Assigned every pattern an owner

    A recurring loss reason became a named fix with a person on it, not a note nobody read twice.

Run the win-loss loop
04 / WHAT I LEARNED

Not proof. Just what the build taught me.

  1. 01
    A model that scores losses like wins is decorating

    Several lost accounts had scored in the 90s on fit. A score that cannot separate the deals you win from the deals you lose is decoration.

  2. 02
    Measure lift, not opinion

    A signal that appears equally in wins and losses is noise, no matter how good it feels. Every signal got compared across won versus lost.

  3. 03
    Keep signals separate so they compound

    One behavioral signal plus an ICP threshold stacked to a 3.56x lift, beating either alone. That is the whole argument for not blending them.

  4. 04
    Retire the weight a signal has not earned

    The fit score and CRM presence both hovered near 1.1x, barely better than a coin flip. They stopped carrying weight.

05 / THE WORKFLOW

The runnable version. Copy it into your stack.

SalesforceAmplitudeClaudeSnowflakeDeeplineSalesforce
[ SALES OPERATOR ]

The Win/Loss Backtest

Heath Barnett · Sales Operator
G / Groundthe problem that started all of it, with the receipt and the cost
Scoring evidence

The model was decorating, not scoring.

Several accounts we lost had scored in the 90s on fit. A model that scores losses as high as wins is not scoring.

Signal

The real predictors were not weighted at all.

The strongest discriminators were behavioral, how the buyer worked, not the firmographics we scored on.

A / Assignthe build, block by block, and who owns each one. Open a step to see it run.
HUMAN + AI, IN THE LOOP
AI measures lift across every signal; the operator decides what earns weight; the pipeline reweights.
Pull the truth
Salesforce01
Pull won and lost deals
Details
The outcomes the model gets tested against, wins and losses side by side.
AI
Amplitude02
Pull the behavioral signals
Details
How the buyer actually worked, sequence activation the loudest.
AI
Measure lift
Claude03
Measure lift per signal
Details
How often each signal shows in wins vs losses. Lift, not vibes.
AI
Snowflake04
Run it on the deal history
Details
The deal-and-signal history the backtest ran across.
AI
Decide the weights
The operator05
Keep what compounds, retire the flatterers
Details
One behavioral signal plus an ICP threshold stacked to 3.56x; the fit score barely beat a coin flip, so it lost its weight.
Human
Deepline06
Reweight the model
Details
The new weights ship into the scoring pipeline the reps run on.
AI + human
N / Normalizethe motion that made it stick
A workflow without a motion is dead. This is how it became the way the team works.
1

A debrief on every closed deal

Wins got read back as honestly as losses, separating the great demo from the warm intro.

2

The finding became a qualification question

The top predictor from the backtest became something reps checked for on live deals, not just a finding in a deck.

3

Every pattern got an owner

A recurring loss reason became a named fix with a person on it, not a note nobody read twice.

T / Tie backthe result it drove, and how they know
RESULT · 01

2.67x lift

top signal: 86% of wins vs 31% of losses.

RESULT · 02

94% vs 62% win rate

Warm-and-above vs the base rate.

RESULT · 03

0 lost deals

ever reached Hot or Strike.

OUTPUT · 01

Per-signal lift table

Wins vs losses, measured.

OUTPUT · 02

The compounding stack

Signals that beat either alone.

OUTPUT · 03

A reweighted model

Weight only where it is earned.

The Sales Operator
06 / HOW YOU DO IT TOO

Here is what I built. Here is how you build it.

The whole thing installs as one plugin. Or grab the skills a la carte. Everything on this page, runnable, on what you paste today.

EVAL PASS · 4/4
THE PLAYBOOK · ONE INSTALL
Turn closed deals into a repeatable edge

4 skills chained into one runnable play. Installs as a single plugin, no copy-pasting each skill. It runs on what you paste; connect your stack to go live.

SEE THE FULL PLAYBOOK →
SEE IT RUN IN CLAUDE
EXAMPLE CHATSam, RevOps Manager, running the loop in one sitting
S
We just closed the quarter. Before everyone forgets, what do our wins actually have in common?
S
Step 1· Closed Won Analysis
I read every won deal and pulled the signals, personas, and motions that keep repeating, then ranked the active accounts that look like them. The pattern is pretty clear.
Receipt
3 of your last 4 wins came in through a product trial, all with a VP champion by week 2.
S
And the ones we lost? I want to know if we're losing on price or something we can fix.
S
Step 2· Closed Lost Analysis
I sorted the losses into no-decision, price, competitor, timing, and fit, then split what's preventable from what's structural. Price is not your problem.
no-decisionhalf
price2 of 11
Receipt
Most losses were no-decision, not lost to a competitor.
S
Can we make this a real program so we're not guessing? Who should we actually talk to?
S
Step 3· Win Loss Program
I built the interview slate, a question guide that won't lead the witness, and a way to code themes across deals so the pattern turns into an actual change with an owner.
Receipt
8 interviews queued, 4 won and 4 lost; messaging fix assigned to a named owner.
S
Pull it all together. What do wins and losses share, and who do we go work now?
S
Step 4· Pattern Analyst
I read wins and losses side by side to find the one thread running through both, then handed back a ranked list of active lookalikes to work.
Receipt
Single-threaded deals stall or lose; 12 active accounts match the winning pattern, ranked.
THE OTHER HALF · LEAD THE TEAM
Run the win-loss loop
I write up one of these a week. Free, receipts only.