B2B account scoring: what actually works
by Morgan
I think most companies get an account score in roughly the same way. Someone says: we have too many accounts. How do we know which ones sales should work?
So the team makes a list. Company size, industry, CRM, website traffic. They give each thing points. And that's the score.
The part that gets missed is checking if any of it works.
Not just if the rules make sense, because they usually do when you read them. The real question is whether the companies at the top actually bought more often than the companies at the bottom.
We learned this at Chili Piper in a pretty painful way.
We spent two years on our account-scoring model. We had 26 SDRs using it to decide where to spend their time. And then we paid an expensive data science consultant to tell us if it worked.
It didn't.
Basically, we could have used a random monkey throwing darts and ended up in the same place.
So, not great.
We rebuilt the whole thing. That story gets quite detailed, so we put it in a separate article about how Chili Piper scores accounts.
What is account scoring?
B2B account scoring is how you rank the companies you could sell to.
You need some way to do this because your team can't work every account. If a rep can properly work 50 accounts this week and there are 100,000 companies in your market, the score decides who gets attention and who doesn't.
There are three different questions hiding inside most account scores:
- Is this company a good fit?
- How valuable could they be as a customer?
- Is there any sign they want to buy now?
Keep those separate.
A company can be a great fit but not be buying this year. Another company can visit your pricing page ten times and still be a customer you don't want.
And lead scoring is not the same thing. Lead scoring ranks people, usually after they do something. Account scoring looks at the company, often before anyone there has talked to you.
Five ways to score accounts
1. Give things points
This is the most common approach.
You pick the things that make a company look like a good fit and give each one points. Maybe Salesforce is worth 50. The right industry is worth 30. A certain amount of website traffic is worth 20.
Starting here is fine, especially if you don't have much data yet.
The problem is that the numbers are guesses. Why is using Salesforce worth 50 points instead of 20? Usually because someone on the team decided it was important. The only way to know if that guess was right is to compare the score with the accounts that actually bought.
So keep the model simple and save every score. Six months or a year later, look at who bought.
If the high-scoring accounts didn't buy more often, change the model.
It sounds obvious. A lot of teams never do it.
2. Use your past deals
If you have enough past deals, you can stop guessing the points and build a propensity model.
You give it examples of companies that bought and companies that didn't. The model looks for what was different between the two groups.
There is one easy way to mess this up: giving the model information from after the sales process started.
Of course a company with an open opportunity, six booked meetings, and a conversation with your VP of Sales looks likely to buy. You don't need a model to tell you that.
Only give the model information you had before sales reached out. Then use old deals to see if its predictions were right.
And keep the original prediction. If yesterday's score disappears when today's score arrives, you have nothing to check later.
3. Look at who stays
Getting someone to buy is not the same as getting a good customer.
Some customers stay for years and grow. Some leave after a few months. A model that only predicts who will sign treats them as the same.
If you have enough customer history, survival analysis can help estimate how long a company might stay. That gives you a better idea of lifetime value, or LTV.
Now you can look at two things: how likely the company is to buy and what the customer could be worth.
You need a decent number of customers for this. If you only have a few, I wouldn't trust what the model tells you.
4. Use AI to find signals
AI is useful for information that doesn't fit nicely into a spreadsheet.
It can read a careers page and pull out what types of roles a company is hiring for. It can read a website and find clues about the business.
Use it to find possible signals. Don't let it make up the final score and trust the answer because it sounds confident.
We all know by now that AI can be very sure and very wrong at the same time.
5. Add intent
Fit and intent answer different questions.
Fit is: do we want this company as a customer?
Intent is: is there any sign they might be buying now?
Use fit to choose the accounts your team cares about. Then use intent to decide who gets contacted first.
Three people visiting your pricing page doesn't turn a bad-fit company into a good one.
What HubSpot, Pardot, and Marketo can do
Most teams start with the tool they already pay for. Fair enough. Here is what each one can actually help with.
HubSpot lets you add and subtract points for things like job title, company size, and website activity. Higher tiers also have predictive scoring. It's easy to set up. Just remember that if your team chose the points, they are still guesses until you test them. Contributed by Costas.
Pardot separates activity and fit. Score tracks what someone does. Grade, from A to F, tracks how well they match the profile your team created. I like the split because clicking every email does not make someone a good fit. But the team still chooses the grading rules, and Pardot is mainly scoring people, not companies. Contributed by Maria.
Marketo lets you build detailed rules around who someone is and what they do. It can reduce a score when activity gets old and send someone to sales when they reach an MQL threshold. It can do a lot. It can also become a lot to maintain. Contributed by Gaines.
All three are better at scoring people who are already engaging with you than deciding which untouched companies sales should go after.
If you want to buy instead of build
You can also pay someone else to do part of this. That makes sense if nobody on your team can own the model.
| Vendor | What it does | Worth a look if... |
|---|---|---|
| GoodFit | Custom company data and fit scores | Teams that need signals missing from standard databases |
| Keyplay (Inflection.io) | Transparent ICP scores with custom signals | SaaS teams that want to see why an account got its score |
| SalesIntel | Company data, verified contacts, and ICP modeling | Teams that want data and scoring from one place |
| MadKudu | Scoring based on conversion, product, and behavior data | PLG teams with enough product data to use it well |
| 6sense | Fit, intent, and predicted buying stage | Enterprise teams running ABM |
| Clay | Data enrichment and scoring logic you build yourself | GTM teams that want control and will maintain it |
| HG Insights / BuiltWith | Data on which technologies companies use | Teams whose ICP depends heavily on tech stack |
A vendor can give you company data and help you score fit. What it doesn't have is your full customer history. It doesn't know who renewed, who expanded, or who left.
That part still has to come from you.
Which method should you choose?
I would start with the simplest thing your data can support.
If you only have 30 or 40 completed sales opportunities, don't try to build a fancy model yet. Use a small point system, write down why each rule is there, and save the scores.
With hundreds of completed opportunities, a propensity model becomes more useful.
If you also have a few hundred customers and know who stayed, you can start looking at retention and LTV.
And if nobody has time to maintain any of this, buy it. A simple model that someone checks is better than a sophisticated one everyone forgot about three months after launch.
Before you use the score
Ask:
- What decision is this score helping us make?
- Are fit, value, and intent separate?
- Did we only use information we had at the time?
- Does the team know what to do with each score?
- Did we save the original predictions?
- Did the high-scoring accounts actually perform better?
- Is anything in the model just because it sounds right?
Number six is the one I care about most.
If nobody knows the answer, stop discussing whether a signal should be worth 20 points or 30 and run the test first.
Frequently asked questions
How often should account scores change?
Fit doesn't change every morning, so updating it monthly is usually enough. Intent can change overnight, so update that weekly or in real time. Then check every quarter if the score is still getting it right.
What if we don't have enough data yet?
Start with a few simple rules and admit that they are guesses. Save today's scores. Later, you can compare them with who bought and build something better.
Do we need a GINI coefficient?
No. But you do need some way to check the score. GINI tells you how much better your list is than a random one. If you don't use GINI, at least check whether the accounts at the top bought more often than the ones at the bottom.
Should we buy a tool or build the model?
Build it only if you have good historical data and someone who will keep checking it. Otherwise, buy. Either way, save the predictions and see if they were right.
Where I would start
I wouldn't start by rebuilding anything.
Pull the scores your team was using before the deals started. Compare the companies at the top with the companies at the bottom.
Who bought? Who stayed? Who left?
Maybe the score works. Maybe it doesn't.
Find out before asking the team to spend another year using it.