Here is one I actually shipped. What happened, what I did, and the stack if you want to run it yourself.
Every scoring model is a set of assumptions until you test it against reality. Ours assumed fit predicted wins.
But when I pulled the won and lost deals and measured what actually separated them, the fit score barely moved the needle, and several accounts we lost had scored in the 90s on fit. A model that scores your losses as highly as your wins is not scoring. It is decorating.
- GThe scoring model assumed fit predicted wins, and several lost accounts had scored in the 90s: decorating, not scoring.
- ISalesforce supplied the won and lost deals, Amplitude the behavioral signals, Claude measured lift per signal, Snowflake held the history.
- AAI measures lift across every signal, the operator decides what earns weight, and the pipeline reweights through Deepline.
- NEvery closed deal, won or lost, got a debrief on the five PLAN questions, and the top predictor became a qualification question.
- TThe strongest behavioral signal carried a 2.67x lift, 3.56x stacked with an ICP threshold, while the trusted fit score sat at 1.13x.