The previous article drew the line between AI for your people and AI in your products. This one is about what happens to that line when it reaches the finance system, because the two sides bill on completely different principles.
One counts people. The other counts work. That difference is why one of them sits quietly in a forecast for a year and the other one produces a phone call.
Per-user licensing: a fixed cost per head
This is the model most organisations already understand, because it is how they buy nearly all their other software.
You pay a fixed monthly fee for each person who has access. If the fee is R400 a month and you license 500 people, you pay R200,000 a month. If those 500 people have a spectacularly productive month and use the tool for six hours a day each, you pay R200,000. If half of them forget it exists, you also pay R200,000.
The strengths are obvious. It is predictable — it can be forecast a year out with confidence, because it moves only when headcount moves. It is familiar — it slots into existing software asset management, existing renewal cycles, existing approval routes. And it is capped — there is no scenario in which enthusiastic usage produces an invoice nobody expected.
The weakness is on the other side of exactly the same property: you pay for the people, not for the value. A licence assigned to somebody who opens it twice a quarter costs the same as one assigned to somebody who has restructured their working week around it. That is the quiet waste in nearly every seat-based estate, and it is invisible unless someone specifically looks at assignment against actual usage. It is also, happily, the easiest AI money any organisation can save.
Consumption pricing: a variable cost per use
The other model bills you for what the AI actually processes.
Nobody is licensed. There are no seats. There is a service running, and every time it does a piece of work — reads a question, considers some of your content, writes an answer — that work is measured and charged. A quiet month is cheap. A busy month is not.
This is the model that feels endless to leaders, and I understand why. There is no natural stopping point built into it, no moment where the thing is bought and paid for. It just keeps metering, the way the electricity meter in your building keeps metering.
But "endless" is the wrong word for it. The right word is proportional. The bill grows because the usage grows, and usage growing is, in almost every case, the thing you were hoping would happen.
Tokens, in plain language
The unit these services meter is the token, and it is worth thirty seconds of explanation because it turns up on every invoice and in every estimate.
A token is roughly a word — a bit less, actually; long or unusual words get split into two or three, and punctuation counts. For practical purposes, treat it as the words going in and the words coming out.
Three things get counted in a single interaction, and only the first is obvious.
What the user typed. The question itself. Usually short.
What you sent along with it. This is the part that surprises people. To answer well, the service typically attaches relevant material — the policy extract, the product details, the customer's own record, the instructions telling it how to behave. That supporting content is often far larger than the question, and it is counted too.
What the AI wrote back. The answer. Usually charged at a higher rate per token than the input, because generating is more expensive than reading.
Add those up and one interaction has a cost. It is a small cost. The reason the monthly figure is not small is arithmetic: a small cost multiplied by a number that grows every time the service gets more popular.
Why this matters more than it sounds
Understanding the unit changes what you can ask for, and that is the practical payoff.
It means you can ask what one interaction costs — a number your team can actually produce, and one that turns an abstract worry into a line of arithmetic you can do in your head. If a resolved query costs a few Rand and it replaces a call that costs considerably more to handle, you have a business case. If it costs more than the call, you have a problem worth naming early.
It means you can ask what is in the payload — because the supporting content attached to each question is a design decision, not a fact of nature. Sending an entire policy manual with every query, when a relevant page would do, is a real and common way to multiply a bill several times over for no gain in answer quality.
And it means you can ask what happens at ten times the volume, which is the question that separates a business case from a hope.
The budgeting mistake to avoid
Here is the thing that actually goes wrong, and it is a filing error rather than a technology one.
The two costs get merged into a single "AI" line in the budget. A number is agreed. Then the consumption side grows — because the service launched, or was promoted, or a competitor's phone line got worse — and it eats the room that had been assumed for the licence side. Or the reverse: a licence renewal lands, absorbs the line, and the customer service has to be throttled to fit.
They do not belong together. One is a fixed cost driven by headcount, reviewed annually, owned by whoever owns software licensing. The other is a variable cost driven by demand, reviewed monthly, owned by whoever owns the service it powers. Merging them means neither one is actually being governed, and the first symptom of that is a surprise.
What to do about it
Three things, none of them difficult.
Split the line. Licences and consumption in separate budget lines, with separate owners, from the beginning. Retrofitting this after a surprising invoice is possible but the conversation is worse.
Ask for cost per interaction, not just total spend. Total spend rising is not information. Total spend rising while cost per interaction falls is a service getting more popular and more efficient. Total spend rising while cost per interaction also rises means something has changed in the design, and you want to know what.
Review licence assignment quarterly. Seats issued versus seats actually used, and reclaim the difference. It is the least interesting recommendation in this series and reliably the fastest saving.
How CloudNala can help
Where we tend to be useful is in producing the per-interaction number before a service launches rather than after — modelling what a realistic month looks like at expected volume, and being explicit about what happens at three and ten times that. It is not a difficult calculation. It is simply one that very few business cases contain, which is why so many of them are approved on a number nobody can defend six months later.
Work with CloudNala
CloudNala helps organisations move from technology ambition to practical execution across cloud, AI, data, platform engineering and digital services.
Whether you are exploring AI, modernising your cloud environment, building a public-sector digital service, or turning an idea into a working MVP, we can help you shape the roadmap and deliver the next step.
Book an AI Readiness Workshop or write to us at consult@cloudnala.co.za