Score every AI use case on two axes: the value it produces and the complexity of building it. Use a narrow 1 to 3 scale on both, score in a group rather than alone, and build only what lands high on value and low on complexity. Everything else waits, with the reason written down.
That sounds mechanical, and it is meant to. The alternative is choosing whichever use case has the most enthusiasm behind it, which is how most AI projects get selected and is a large part of why they fail. RAND interviewed sixty-five experienced data scientists and engineers for its study on the root causes of failure for AI projects and found that more than eighty percent fail, roughly twice the rate of IT projects without AI. The first root cause on their list is not a technical one. It is that stakeholders misunderstand or miscommunicate which problem is being solved.
Scoring is the cheapest available fix for that. It forces the problem to be stated in numbers before anyone commits budget, and it produces a written record of why the other thirty ideas were set aside. This article covers what scoring actually involves, how to calculate each axis, where the thresholds should sit and which mistakes make the exercise worthless.
What does scoring an AI use case actually mean?
It means assigning two numbers to a described process, not to an idea. That distinction does most of the work. “Use AI in customer service” cannot be scored, because it is a direction rather than a task. “Classify inbound service emails by product line and urgency, currently done by two people for roughly ninety minutes a day” can be scored, because every part of it is countable.
So the first output of any scoring exercise is a rewritten list. Ideas come out of interviews as intentions, and each one has to be converted into a process description with a volume, a frequency and a current owner before a score means anything. In practice this step removes ten to twenty percent of a longlist on its own, because some ideas turn out to describe the same process twice and others turn out to describe no process at all.
The second thing worth being clear about is what scoring is not. It is not a business case, which is a fuller calculation of hours, build cost and payback for a single use case. It is not a feasibility study either. Scoring is deliberately shallow and fast, perhaps five minutes per use case, precisely so that the expensive analysis only gets spent on the few that survive.
Why does scoring matter more than the idea list?
Because a long list without a ranking is as paralysing as no list at all. In the assessments we run, an organisation of five departments typically produces between seventy and a hundred and forty distinct opportunities. Handed over unranked, that document gets read once and then quietly stops being used, because there is no defensible way to start.
The failure pattern that follows is well documented. Gartner expects that more than forty percent of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. Unclear business value is a scoring failure by another name. If nobody put a number on the value before the build started, nobody can defend the spend when it doubles.
There is a second, less obvious argument for scoring. It creates an audit trail. Six months into a programme, somebody will ask why their department’s request never made the roadmap, and the honest answer needs to be more than a shrug. A score with a one-line justification answers that question in a way that keeps people engaged rather than resentful.
How do you score value without guessing?
Value has to be a number you can recompute, which rules out most of what gets written on sticky notes. Start with the time the process consumes today, then add the cost of the errors it currently produces, and only then consider revenue.
The time calculation is straightforward once the process description is precise. Take the hours spent per occurrence, multiply by frequency across a year, and multiply again by a fully loaded hourly cost rather than a salary figure. The number that catches people out is the last multiplier: the share of the task that automation genuinely removes. That is rarely the whole thing. In document-heavy processes we typically see forty to seventy percent removed, with the remainder becoming review and exception handling rather than disappearing. Scoring a use case as though the task vanishes entirely is the single most common way value gets overstated.
Error cost is the axis most teams skip, and it is often larger than the time saving. A process that runs at a two percent error rate across twenty thousand transactions produces four hundred errors a year, and each one carries a rework cost, sometimes a credit note, occasionally a customer. Multiply and compare. Where a process touches compliance or safety, the error cost dominates the calculation entirely and the time saving becomes incidental.
Revenue effects are real but should be counted conservatively, and only when someone will own the number. A faster quotation process plausibly wins more work, but unless sales can state a current conversion rate and commit to measuring it afterwards, treat it as a qualitative note rather than a score component. Optimistic revenue attributions are how scoring exercises lose credibility with a finance director.
With those three inputs, the value score becomes simple. A use case scores 1 when the annual value sits below roughly two hundred hours or cannot be quantified at all, 2 when it lands between two hundred and a thousand hours or carries a documented error cost, and 3 when it exceeds a thousand hours or has a revenue effect with a named owner. Adjust those thresholds to your own scale, but fix them before you start scoring rather than during.
How do you score complexity?
Complexity is the axis where technical teams and business teams disagree most, so it helps to break it into five factors and score the worst one rather than the average.
Data availability comes first, because it is the factor that most often turns a promising use case into a six-month project. The question is not whether the data exists, but whether it exists in a form a system can read, in a place a system can reach, with a history long enough to test against. Scanned documents in a shared folder and structured records in a database sit at opposite ends of this scale.
Integration count is the next factor, and it scales worse than people expect. One system is straightforward. Two systems with a documented API is manageable. Three systems where one has no API at all is a different project entirely, and belongs at the top of the complexity scale regardless of how simple the AI component looks.
Exception rate deserves more weight than it usually gets. A process where ninety-five percent of cases look alike is a good automation candidate. A process where every third case has something unusual about it will need exception handling that costs more to build than the happy path. Ask the person who does the work how often they have to think, rather than asking their manager how standard the process is.
Regulatory class matters in the EU, because the AI Act’s classification of high-risk systems determines what documentation, oversight and testing you are obliged to put in place. A use case that touches recruitment, credit, education or essential services carries obligations that materially change the build effort, and that belongs in the complexity score rather than being discovered afterwards.
Behavioural change is the factor most often left out, and it is the reason some technically trivial use cases never land. If the output only creates value when forty people change how they start their day, the complexity is high even though the engineering is not. Count the number of people whose habits have to change, and treat anything above a single team as a complexity driver.
Score 1 when the data sits in one system, no new integrations are needed and the exception rate is low. Score 2 when two or three systems are involved and some data preparation is required. Score 3 when the data does not yet exist in usable form, when an integration has no API, when exceptions are frequent, or when the use case falls into a high-risk regulatory class.
What do you do with the two scores?
Plot them and cut. The two axes produce four quadrants, and only one of them gets built in the coming quarter.
High value with low complexity is where you start, and there are usually fewer of these than anyone hopes. High value with high complexity is not a rejection but an instruction: break it into a smaller first version that can be built and measured, then rescore the remainder. Low value with low complexity is genuinely tempting and worth resisting, because a portfolio of small easy wins consumes the same management attention as one significant project while producing far less. Low value with high complexity gets parked with a note.
The narrow scale is deliberate. Scoring on 1 to 10 feels more precise and produces worse decisions, because most scores land between four and seven and the ranking collapses into noise. Three points force every use case to be called low, medium or high, and disagreement about which one becomes a useful conversation rather than an averaging exercise.
One rule keeps the cut honest: write a single sentence next to every use case that does not survive, saying why. That sentence takes fifteen seconds and it is what makes the exercise defensible later.
Which scoring mistakes make the whole exercise useless?
Five patterns account for most of the damage, and all five are avoidable.
- Scoring the idea rather than the process. If the description does not contain a volume, a frequency and a name, there is nothing to score. Rewrite it first.
- Scoring value on strategic importance. Strategic matters, but it is not a number, and once it enters the scoring column every executive sponsor’s project becomes a 3. Keep strategic weight as a separate tiebreaker applied after scoring, not inside it.
- Averaging the complexity factors. A use case with perfect data and one unreachable system is not medium complexity. It is blocked. Score the worst factor.
- Scoring alone. One person produces confident numbers with no shared basis. Three to five people, including someone who actually runs the process, produce numbers that survive scrutiny.
- Treating a high score as permission to build. Scoring ranks candidates. It does not confirm that the data is accessible, that an owner has time allocated or that success can be measured. That is a separate gate, and skipping it is where scored roadmaps still fall apart.
How do you keep the scores honest six months later?
Rescore rather than reopen. Circumstances change, particularly on the complexity axis: a system migration completes, an API becomes available, a data quality project finishes, and a use case that scored 3 for complexity legitimately becomes a 2. That is a reason to rescore it in the next cycle, not to argue it back onto the current roadmap.
Set a fixed rhythm for that, quarterly in most organisations, and keep the parked list visible in the meantime. A parked use case with a stated reason and a review date behaves very differently inside an organisation than a rejected one. People keep contributing ideas when they can see where their previous ones went.
The other half of staying honest is checking your own estimates against what actually happened. Once the first use case reaches production, compare the value you scored against the value you measured. Most teams discover their time savings were optimistic and their error costs were understated, and both corrections make the next round of scoring better.
What scoring does not tell you
A score ranks candidates against each other. It does not tell you whether the top candidate is ready to be built, and those are genuinely different questions. A use case can score 3 on value and 1 on complexity and still stall, because the data turned out to be locked in a system nobody has credentials for, or because the business owner has no time allocated, or because nobody agreed what a good result would look like.
That is why scoring belongs directly in front of a readiness check rather than in front of a development sprint. Rank first, then verify. Doing it in that order costs a fortnight and saves the six months that a well-chosen but unready use case tends to consume.
Frequently asked questions about scoring AI use cases
What scale should you use to score AI use cases?
Use 1 to 3 on both value and complexity. Wider scales such as 1 to 10 invite hedging, and scores cluster in the middle where they carry no information. Three points force a decision on every use case, which is the entire purpose of scoring.
How do you calculate the value of an AI use case?
Multiply the hours currently spent on the process by the frequency and the fully loaded hourly cost, then add the cost of errors the process produces today. Count only the share automation actually removes, which is rarely the whole task and often between 40 and 70 percent.
How many use cases should survive scoring?
Three to five, out of the seventy to a hundred and forty a typical inventory produces. If more than five survive, the scoring was too generous. A roadmap carrying ten use cases usually means nobody was prepared to make a choice.
Does scoring replace a business case?
No. Scoring decides which use cases deserve a business case, which is a more detailed calculation of hours, build cost and payback. Scoring is deliberately cheap and fast so that you only spend real analysis time on the handful that survive.
Who should score the use cases?
Score in a group of three to five people that includes someone who runs the process, someone who knows the systems and someone with budget authority. Individual scoring produces confident numbers that fall apart the moment anyone asks how they were reached.