Guide
AI lead scoring: how to make a score you can actually defend
Most lead scores are a number with nothing behind them. This is how signal-based scoring works, what it can read off public sources, and how to write an ICP that produces a list worth calling.
1. What lead scoring is for
Lead scoring exists to answer one question: of the names in front of me, which ones do I call first? That is it. A score is not a prediction of revenue and not a measure of company quality. It is a ranking device for your own limited hours.
That framing matters because it sets the bar for a good score. A good score is one that changes the order you work the list. If every lead scores between 70 and 80, the model has told you nothing, no matter how confident the number looks.
2. Why most scores are useless
Three failure modes show up over and over.
- The black box. A number appears with no reasoning. You cannot tell whether it is reading something real or pattern-matching on company size, so you either trust it blindly or ignore it.
- Missing data treated as bad data. A company with no website scores low. But no website is sometimes the strongest buying signal in the set, depending on what you sell. Absence is not a negative unless you decided it is.
- Firmographics dressed up as intent. Industry, headcount, and revenue band describe who a company is, not whether anything has changed for them lately. Fit is not intent.
3. Signals, tiers, and weights
A signal is a checkable statement about a business. Not "is a good fit" but "the About page names the owner" or "a review in the last 60 days mentions an unanswered call." Checkable means a person could go verify it in ten seconds.
Signals need tiers, because they are not equal. A workable split is two levels:
- A-tier signals define the fit. If none of them hit, this is not your customer. Weight them 3.
- B-tier signals are supporting evidence. They break ties between leads that already clear the A-tier bar. Weight them 1.
Then every signal gets three possible verdicts, not two: hit, miss, and unknown. Unknown is the one most systems skip, and skipping it is what produces confidently wrong scores. If the source needed to check a signal was not available, the honest answer is unknown, and unknown must contribute zero in either direction.
4. Anti-signals
Positive signals rank the list. Anti-signals remove names from it. They are the disqualifiers you would state out loud: already using a competitor, too large to care, wrong geography, franchise with no local decision-maker.
Anti-signals should not shade a score down by a few points. One clear hit should zero the lead, because a disqualified prospect with six positive signals is still disqualified. Shading produces the worst possible outcome: a well-evidenced lead near the top of the list that cannot buy.
5. Worked example
Real numbers from a pull of HVAC contractors in Nashville, scoring for call-handling software. Five A-tier signals, five B-tier. A hits count 3, B hits count 1.
| Signal | Tier | Verdict | Evidence |
|---|---|---|---|
| Missed-call complaints in reviews | A | HIT | "Called three times before anyone picked up" |
| Review count in the 100–2,000 band | A | HIT | 142 ratings |
| About page names the owner | A | HIT | "Family owned by the Delgado family since 1998" |
| Not a national franchise name | A | HIT | No exclusion-list match |
| Advertises 24/7 or after-hours service | A | MISS | "Mon–Fri 7am–5pm" in the footer |
| Fresh review in the last 30 days | B | HIT | Most recent review 2 days ago |
| Under 10 employees | B | HIT | "our six technicians" |
| Single phone number on the site | B | MISS | Three numbers listed |
| Access friction in reviews | B | MISS | Nothing beyond the review above |
| Profile photos show a truck | B | UNKNOWN | No photos returned for this listing |
| 4 A-tier hits × 3 + 2 B-tier hits × 1 | funnel 14 → HOT | ||
Note what the unknown row did: nothing. It did not penalise the lead for a gap in Google's response, and it did not get quietly counted as a miss to make the math look complete.
Then set a bar. Two thresholds are enough: a minimum number of A-tier hits, and a minimum funnel score. A lead under either one is dropped, and the drop reason is recorded so you can tell whether your bar is too high.
| Funnel score | Tier | What to do |
|---|---|---|
| 9 or more | HOT | Call this week |
| 6 to 8 | WARM | Worth a sequence |
| 3 to 5 | COLD | Only if your A-tier bar is 1 |
| Below bar | DROP | Keep the reason, do not delete the row |
6. Writing the ICP
The ICP is the input that decides everything downstream. Write it as three blocks.
- Who, in one sentence. "Owner-operator HVAC companies in Middle Tennessee doing their own phone handling." Specific enough that a stranger could sort a list with it.
- Positive signals, as checkable statements. Not "established business" but "100 to 2,000 Google reviews." Not "growing" but "a review in the last 30 days."
- Disqualifiers, stated flatly. "Already running CallRail or WhatConverts." "Three or more locations." These become your anti-signals.
Then tune on one query before you run a hundred. Pick a market you know by eye, run twenty leads, and read the evidence quotes on the ones you disagree with. Nine times out of ten a bad score traces to a vague signal, not a bad model.
7. What public data cannot tell you
Public sources are enough to rank a list. They are not enough to know a business. Here is the honest boundary.
| Readable today | Not readable from public sources |
|---|---|
| Review text, count, rating, recency | Contract renewal dates |
| Listing category, address, phone, website | Budget or revenue |
| Website copy, vendor scripts, hours, team size claims | Who actually signs |
| Whether a competitor's tooling is installed | Internal dissatisfaction that was never written down |
Hiring posts, ownership-change records, and trade-directory membership are all readable in principle, but each needs its own source wired up. Until that work is done, the correct behaviour is to report unknown, not to infer.
8. Five common mistakes
- Scoring before filtering. Run cheap mechanical filters first: no website, no phone, PO box addresses, thin review counts. Never spend a model call on a lead a regex could have rejected.
- Too many A-tier signals. More than about five and everything scores HOT. Tiering only works if A-tier means something.
- Weights with decimals. If you are arguing about 2.5 versus 2.75, the signal set is the problem, not the weight.
- Deleting dropped leads. Keep them with their drop reason. That list is the only evidence of whether your bar is calibrated.
- Never rereading the evidence. Scores drift as your market changes. Read ten evidence quotes a month and you will catch it early.
9. Questions
How many signals should a profile have?
Eight to twelve total, with no more than five at A-tier. Past that, signals start overlapping and the score stops discriminating.
Should the model score, or should rules score?
Both. Rules do the mechanical checks: review counts, banned keywords, address shape. The model does the reading: what a review actually complains about, whether an About page names a human. Asking a model to do arithmetic it cannot audit is how you get soft numbers.
How often should I rewrite the ICP?
When you notice yourself skipping leads the model ranked highly. That gap is the signal that your real criteria have moved and the written ones have not.
Run this on your own market.
14-day trial, full Solo access, no credit card. One query on a market you know is enough to tell whether the scoring is reading anything real.