How to tell whether an AI project paid for itself
Almost every AI business case is built after the fact, from an estimate of what things used to take. The measurement that would have settled it costs an afternoon and cannot be done retroactively.
Ask a company whether its AI spend has paid back and you get an anecdote. Not because nobody is trying to be rigorous, but because the number that would settle it, how long the work took before, was never recorded and cannot be recovered once the process has changed.
Three numbers, recorded before anything changes
Time. How long does the task take today, measured rather than estimated? An estimate of a task someone does fifty times a week is a guess about a habit, and the error gets multiplied by the frequency below.
Frequency. How many times a month does it run? This is the multiplier, and it decides most cases on its own. A two-hour task done monthly is twenty-four hours a year. A four-minute task done six hundred times a month is forty hours a month.
Quality floor. What does an unacceptable output look like, written down before you have seen the system's output? Defining it afterwards means defining it around what the system happens to produce.
The costs that get left out
Business cases routinely count the licence and stop.
- Review time. If a person checks the output, that is part of the running cost forever, not a temporary phase. Frequently it is the largest line.
- The wrong answers that get through. Some proportion will not be caught. What does one cost? For internal drafting, very little. For anything that reaches a customer or a regulator, potentially more than the whole saving.
- Maintenance. Integrations break, providers retire model versions, processes change. Budget for the second year, which most proposals do not.
- The change itself. Time spent getting people to actually use it, which is real and is usually borne by whoever least wanted the project.
Measure the thing, not the sentiment
Adoption surveys tell you whether people liked it. That is worth knowing and it is not the business case. Two things are worth instrumenting from day one:
- Usage against the population. How many of the people who could use it did, in the last thirty days? Where that sits below the licence count you have the same problem as seats nobody sits in, on a newer invoice.
- The before-and-after on the three numbers. Same task, same conditions, measured the same way.
Where the honest answer is that it did not pay back, say so and stop. The main cost of an unmeasured AI portfolio is not the tools that failed. It is that nobody can tell which ones failed, so all of them get renewed.
What to do this week
Pick the AI project you are most likely to fund next quarter. Before it starts, spend one afternoon timing the task it targets and counting how often it runs. Write both numbers down somewhere durable with today's date on them.
That afternoon is the entire difference between a business case and an argument, and it is only available before the project starts.
Where this goes next
Frequency and specificity also settle whether to build at all: build an AI agent or buy an AI tool. The reason projects reach this question having skipped the measurement is in why AI pilots do not reach production.
Sizing the payback before anyone builds, and saying when it is not there, is part of AI consulting.