Guide

AI lead scoring: how to make a score you can actually defend

Most lead scores are a number with nothing behind them. This is how signal-based scoring works, what it can read off public sources, and how to write an ICP that produces a list worth calling.

1. What lead scoring is for

Lead scoring exists to answer one question: of the names in front of me, which ones do I call first? That is it. A score is not a prediction of revenue and not a measure of company quality. It is a ranking device for your own limited hours.

That framing matters because it sets the bar for a good score. A good score is one that changes the order you work the list. If every lead scores between 70 and 80, the model has told you nothing, no matter how confident the number looks.

2. Why most scores are useless

Three failure modes show up over and over.

  • The black box. A number appears with no reasoning. You cannot tell whether it is reading something real or pattern-matching on company size, so you either trust it blindly or ignore it.
  • Missing data treated as bad data. A company with no website scores low. But no website is sometimes the strongest buying signal in the set, depending on what you sell. Absence is not a negative unless you decided it is.
  • Firmographics dressed up as intent. Industry, headcount, and revenue band describe who a company is, not whether anything has changed for them lately. Fit is not intent.
The test. Ask a scoring tool why a lead scored what it scored. If the answer is not a quote from a specific piece of text, the score is a vibe.

3. Signals, tiers, and weights

A signal is a checkable statement about a business. Not "is a good fit" but "the About page names the owner" or "a review in the last 60 days mentions an unanswered call." Checkable means a person could go verify it in ten seconds.

Signals need tiers, because they are not equal. A workable split is two levels:

  • A-tier signals define the fit. If none of them hit, this is not your customer. Weight them 3.
  • B-tier signals are supporting evidence. They break ties between leads that already clear the A-tier bar. Weight them 1.

Then every signal gets three possible verdicts, not two: hit, miss, and unknown. Unknown is the one most systems skip, and skipping it is what produces confidently wrong scores. If the source needed to check a signal was not available, the honest answer is unknown, and unknown must contribute zero in either direction.

4. Anti-signals

Positive signals rank the list. Anti-signals remove names from it. They are the disqualifiers you would state out loud: already using a competitor, too large to care, wrong geography, franchise with no local decision-maker.

Anti-signals should not shade a score down by a few points. One clear hit should zero the lead, because a disqualified prospect with six positive signals is still disqualified. Shading produces the worst possible outcome: a well-evidenced lead near the top of the list that cannot buy.

5. Worked example

Real numbers from a pull of HVAC contractors in Nashville, scoring for call-handling software. Five A-tier signals, five B-tier. A hits count 3, B hits count 1.

SignalTierVerdictEvidence
Missed-call complaints in reviewsAHIT"Called three times before anyone picked up"
Review count in the 100–2,000 bandAHIT142 ratings
About page names the ownerAHIT"Family owned by the Delgado family since 1998"
Not a national franchise nameAHITNo exclusion-list match
Advertises 24/7 or after-hours serviceAMISS"Mon–Fri 7am–5pm" in the footer
Fresh review in the last 30 daysBHITMost recent review 2 days ago
Under 10 employeesBHIT"our six technicians"
Single phone number on the siteBMISSThree numbers listed
Access friction in reviewsBMISSNothing beyond the review above
Profile photos show a truckBUNKNOWNNo photos returned for this listing
4 A-tier hits × 3 + 2 B-tier hits × 1funnel 14 → HOT

Note what the unknown row did: nothing. It did not penalise the lead for a gap in Google's response, and it did not get quietly counted as a miss to make the math look complete.

Then set a bar. Two thresholds are enough: a minimum number of A-tier hits, and a minimum funnel score. A lead under either one is dropped, and the drop reason is recorded so you can tell whether your bar is too high.

Funnel scoreTierWhat to do
9 or moreHOTCall this week
6 to 8WARMWorth a sequence
3 to 5COLDOnly if your A-tier bar is 1
Below barDROPKeep the reason, do not delete the row

6. Writing the ICP

The ICP is the input that decides everything downstream. Write it as three blocks.

  1. Who, in one sentence. "Owner-operator HVAC companies in Middle Tennessee doing their own phone handling." Specific enough that a stranger could sort a list with it.
  2. Positive signals, as checkable statements. Not "established business" but "100 to 2,000 Google reviews." Not "growing" but "a review in the last 30 days."
  3. Disqualifiers, stated flatly. "Already running CallRail or WhatConverts." "Three or more locations." These become your anti-signals.

Then tune on one query before you run a hundred. Pick a market you know by eye, run twenty leads, and read the evidence quotes on the ones you disagree with. Nine times out of ten a bad score traces to a vague signal, not a bad model.

7. What public data cannot tell you

Public sources are enough to rank a list. They are not enough to know a business. Here is the honest boundary.

Readable todayNot readable from public sources
Review text, count, rating, recencyContract renewal dates
Listing category, address, phone, websiteBudget or revenue
Website copy, vendor scripts, hours, team size claimsWho actually signs
Whether a competitor's tooling is installedInternal dissatisfaction that was never written down

Hiring posts, ownership-change records, and trade-directory membership are all readable in principle, but each needs its own source wired up. Until that work is done, the correct behaviour is to report unknown, not to infer.

8. Five common mistakes

  1. Scoring before filtering. Run cheap mechanical filters first: no website, no phone, PO box addresses, thin review counts. Never spend a model call on a lead a regex could have rejected.
  2. Too many A-tier signals. More than about five and everything scores HOT. Tiering only works if A-tier means something.
  3. Weights with decimals. If you are arguing about 2.5 versus 2.75, the signal set is the problem, not the weight.
  4. Deleting dropped leads. Keep them with their drop reason. That list is the only evidence of whether your bar is calibrated.
  5. Never rereading the evidence. Scores drift as your market changes. Read ten evidence quotes a month and you will catch it early.

9. Questions

How many signals should a profile have?

Eight to twelve total, with no more than five at A-tier. Past that, signals start overlapping and the score stops discriminating.

Should the model score, or should rules score?

Both. Rules do the mechanical checks: review counts, banned keywords, address shape. The model does the reading: what a review actually complains about, whether an About page names a human. Asking a model to do arithmetic it cannot audit is how you get soft numbers.

How often should I rewrite the ICP?

When you notice yourself skipping leads the model ranked highly. That gap is the signal that your real criteria have moved and the written ones have not.

This is how IntentScoring works. A-tier and B-tier signals, three verdicts, quoted evidence on every one, anti-signals that zero a lead, and a stated bar. $29 a month for 500 scored leads. See the pipeline or start the trial.

Run this on your own market.

14-day trial, full Solo access, no credit card. One query on a market you know is enough to tell whether the scoring is reading anything real.