Everything in this series so far explains why the AI bill behaves the way it does. This article is about keeping it under control, and the short version is that the controls are not the hard part.
The controls are well understood, unremarkable, and mostly borrowed from cloud financial management, which most organisations have already had a go at. What is usually missing is not a technique. It is an owner.
Nobody in the room can answer the question
Here is the structural problem, and it is worth being precise about because it is not anybody's fault.
Finance receives the invoice. They can see the total, and they can see it moving. They cannot tell which service caused which portion of it, or whether the movement was good news.
IT and engineering can see the usage in detail. They know exactly which workload consumed what. They are not, however, accountable for whether that workload was worth running.
The service owner — the person responsible for the customer journey the AI sits inside — knows precisely what it is worth. They usually have no visibility of what it costs, and often no idea that they could.
Three people, three thirds of the picture, each able to defend their own third completely. And no one of them can answer the only question that matters: was that spend worth it?
This is why AI cost governance fails in organisations that are perfectly competent at cost governance generally. The failure is not analytical. It is that the analysis has no home.
The controls, briefly
With an owner in place, the practical controls are short and mostly obvious. They are listed here in the order they should be applied, because applying them in the wrong order is how organisations end up retrofitting limits onto services people already depend on.
Right-size at design time. Choose the model the job needs, not the largest one available. Trim what gets attached to each request. Cache answers to questions that repeat — and in any real service, a large share of questions repeat. All of this is dramatically cheaper to do at design time than to retrofit, because retrofitting means changing something users have already come to rely on.
Launch small and scale on evidence. Start on demand. Start with a subset of users or a single journey. Let the usage pattern reveal itself before committing to anything with a monthly floor.
Instrument before you launch, not after. Cost attributed per service, per journey, and ideally per interaction — visible from day one. A service that has been running for six months without cost attribution cannot be optimised, only guessed at.
Set budgets and alerts. A threshold with an alert attached, agreed before launch, so that an unusual month produces a notification rather than a discovery. This is trivial to configure and startlingly often absent.
Review commitments on a schedule. Any reserved capacity gets a monthly utilisation check for the first year. A commitment nobody is watching is a subscription nobody cancels.
Measure cost per outcome, not just total cost. The number that makes every other number interpretable.
The six numbers
If a leader takes one practical thing from this series, let it be this list. These are the numbers to require on a monthly dashboard — and requiring them consistently will surface every failure mode in this series early enough to act on.
A note on each of the two that get argued about.
Spend split by kind matters because merging licence and consumption costs into one AI total destroys the information in both. They move for different reasons and are managed by different means. One merged number tells you nothing except that it went up.
Cost per outcome is the one people resist, because agreeing the denominator requires a conversation about what the service is actually for. That conversation is the point. An organisation that cannot say what one unit of output from its AI service is worth has not finished designing the service.
Who owns it
Not IT alone. Not finance in the dark. Both, with the service owner in the room.
In practice the arrangement that works is unremarkable: a standing monthly review, half an hour, with a named chair. Finance brings the spend. Engineering brings the usage and the technical explanation for any movement. The service owner brings the outcomes. The dashboard is the same six numbers every month, and somebody is answerable for each of them.
What makes it work is not the meeting. It is that the movement in the bill has to be explained by someone who understands both halves. "Consumption rose eighteen percent" is not an explanation. "Consumption rose eighteen percent because resolved queries rose twenty-two percent, so cost per resolved query fell" is an explanation, and it is also good news that would otherwise have looked like a problem.
What good and bad look like
It is worth being concrete about what the review is looking for, because "the bill went up" is not by itself a finding.
Healthy. Total spend rising while cost per outcome falls or holds. Reserved capacity utilisation high. Licence assignment close to licence usage. Volume growth traceable to something you did on purpose.
Worth investigating. Cost per outcome rising — something changed in the design, or the service is escalating more than it resolves. Reserved capacity utilisation below half — you are paying for a lane you are not driving on. A large gap between licences issued and licences used — the quietest waste in the building and the fastest saving available.
Actually alarming. Spend rising with no corresponding rise in usage at all. That is usually a defect: something retrying in a loop, a runaway job, or a misconfiguration. It is rare, and it is the reason the alert threshold exists.
Scale spend with value, not ahead of it
The single sentence that summarises this article, and arguably the series: let the spending follow the evidence.
Almost every expensive AI mistake in the last two years has been an organisation committing ahead of proof. Buying licences for everyone before knowing whether anyone would use them. Reserving capacity before the traffic existed. Building for a scale that had not arrived. In every case the money went out before the evidence came in, and in every case the evidence, when it arrived, would have suggested something smaller.
The discipline is not to spend less on AI. Organisations that under-invest here will find that out too, more slowly and more painfully. The discipline is to make each increment of spend conditional on the previous increment having produced something.
That is not a technology practice. It is ordinary management applied to an unfamiliar cost shape.
How CloudNala can help
We set up the review rather than run it forever: the six numbers, wired to real data, with an owner named for each and a threshold that alerts before anyone is surprised. The pattern we most often correct is a service six months live with no cost attribution at all — where the only available lever is a blunt one, because nobody can see which part of the bill is which.
Work with CloudNala
CloudNala helps organisations move from technology ambition to practical execution across cloud, AI, data, platform engineering and digital services.
Whether you are exploring AI, modernising your cloud environment, building a public-sector digital service, or turning an idea into a working MVP, we can help you shape the roadmap and deliver the next step.
Request a Cloud Review or write to us at consult@cloudnala.co.za