The reframe: a score is a bet, not a grade
Most lead scoring guides treat the score as a report card you hand the lead, as if a higher number were a compliment. That framing is the original sin, because it quietly recasts a resource-allocation problem as a flattery problem. A score does not exist to tell a lead how impressive they are. It exists to ration the single scarcest thing you own, which is your attention, and to point it at the handful of conversations most likely to turn into revenue.
The single-number trap
The damage starts the moment you blend everything into one figure. Picture a composite of 72. One lead behind that number is a VP of operations at a company that looks exactly like your best client, who landed on your pricing page once and went quiet. Another is a curious student who has read your entire blog, attended a webinar and downloaded three guides. Identical score, opposite reality. One is a high-value bet waiting on the right moment, the other is engagement noise that will never buy. Average them and you have built a machine that confuses activity with value, which is precisely the confusion a scoring model was supposed to remove.
The number hides the only distinction that matters. It cannot tell you whether you are looking at a perfect-fit account that has not moved or a poor-fit account that happens to be active, and so it cannot tell you what to do next. You end up chasing the busy student and letting the quiet VP cool, which is the exact inversion of a model that earns its keep.
Why this bites a solo operator hardest
In a funded team a mis-scored lead is an annoyance, a rep's wasted afternoon absorbed by the org. When you run growth alone with AI agents, you are the whole funnel, and every follow-up you send is an hour you did not spend on the lead that would have closed. The cost of a bad bet is not diffused across a team, it lands entirely on you. That is why low volume is an argument for scoring, not against it: the fewer hours you have, the more it matters that each one goes to the right conversation.
The only test that counts
There is exactly one honest way to judge a model, and it is not the symmetry of your point table or the cleverness of your weightings. It is conversion lift against an unscored baseline. Take the leads the model ranked highest, take a random unscored sample, and compare how each cohort actually converts. A working model lands somewhere between three and ten times the baseline. If the lift sits at one and a half times or lower, the model is barely separating signal from noise, and it is not worth keeping however elegant it looks on the page. Everything that follows in this guide exists to move that one number, and you should refuse to call a model finished until you have measured it.