In almost every office building in the Netherlands, an AI system is now running somewhere in the background. It answers questions, summarises documents, writes reports. Nobody sees a meter running or a light flashing. There's no warehouse running empty. Yet costs rise steadily, invisibly and sometimes exponentially. The invoice only arrives at the end of the month, and then many a finance director is shocked: how can this amount be so high, and where does it come from?
This is the 'token economy'. If you deploy AI without understanding how that economy works, you're paying for something you can't measure, can't control and can't account for. Not because the technology is poor, but because the economy behind it works fundamentally differently from everything we know from traditional software and even from the cloud.
A token is the smallest unit with which an AI language model processes information: a word, part of a word, a number or a symbol. Every question you ask is broken down into tokens. Every answer likewise consists of tokens. AI providers charge based on what you consume, for both input and output. Customers sometimes call them 'credits', 'points' or even 'floeppies'. That confusion about the term says enough: this cost model is still completely new for many organisations.
You can compare it to your mobile data subscription. You choose a smaller or larger package and the price per unit falls as you buy more. But because your needs structurally grow, you still end up paying more overall. The crucial question is therefore not how much you consume, but whether you're consuming what you need and deploying it for the right things.
That's where it starts to chafe. Many organisations work with an 'all-you-can-eat model': unlimited access via one large cloud contract, without insight into what they actually consume. That feels comfortable, but management becomes impossible: it's not unlimited, just too much for one person and too little for another.
That 'flat fee feeling' won't last much longer anyway. Many 'unlimited' AI subscriptions have until now been partly subsidised by the provider, as a growth strategy to gain market share. Now the sector is switching to pure consumption pricing per token, that subsidy is disappearing. For developers this is already noticeable: since April 2026 the flat rate period for AI tools has largely been vanishing, and the question is shifting from 'how much are we spending on AI' to 'is that expenditure delivering measurable output'. For some users the invoice therefore rises to a multiple of the original price per user (the so-called 'seat price'), precisely the risk you overlook with an unmeasured 'all-you-can-eat contract'.
So, AI costs are rising and then we try to control them. The easiest thing is to set a limit. That response is understandable but solves the wrong problem.
Many organisations see the token economy as a cost control problem. We see a governance problem. It's not about how much you spend on AI, but about the question of whether you know what each token delivers. If you miss that distinction and immediately reach for a ceiling, you save in the short term and destroy value in the long term. Cost control is setting and enforcing a ceiling. Governance is understanding what each pound delivers and choosing on that basis where to invest.
Take a large, fast-moving multinational, which is a leader in AI adoption. The culture is entrepreneurial: experiment, act quickly, don't slow down innovation with bureaucracy. In practice that means: employees use the corporate credit card to take out cloud subscriptions themselves and purchase AI tools, without central approval.
The result? Finance sees costs rising month after month, without anyone being able to trace where they come from, who causes them and what they deliver. That creates a profit and loss account without governance.
At the other end of the spectrum: a public institution. Here there's no question of unrestrained growth, but of the opposite problem. The institution must already plan next year's budget with token consumption as a variable, whilst the justification is lacking. How many tokens does automated processing of tax returns consume? What does fraud analysis cost per file? Now that's guesswork.
One organisation spends uncontrollably, the other can't plan. Yet it's the same diagnosis: neither can manage, because neither links consumption to value. This isn't a cost problem. It's a governance problem.
At this point that pragmatic colleague says: models are becoming drastically cheaper, so why invest in expensive governance? Set a ceiling and move on. Two things are wrong with that. Firstly, the price per token is falling, but consumption per task rises explosively as soon as agents take over work processes. Secondly: that 'cap' isn't free.
An example: an organisation sets a limit of twenty per cent per user per month. For most users their budget remains largely unused: pure waste. For heavy users a stream of exception requests arises; they're at their ceiling and submit requests that someone must assess and process. Ten employees then spend half an hour weekly handling these. Token costs fall, but part of the saving disappears in coordination time, whilst both under- and over-use continues to cost value.
Poorly designed governance creates its own inefficiency. In fact: a uniform 'cap' is a panic measure. For one role it means waste; the budget isn't used anyway. For another a productivity brake; the budget is exhausted, whilst value creation continues. You limit everyone equally and manage nobody well. '
The approach that works for frontrunners is differentiated. Approximately eighty per cent of users are automatically guided via routing to the most suitable model. They choose nothing; the system determines based on the task which model is most cost-effective. A simple classification goes to a small, cheap model; a complex analysis to a heavier one. For this group governance is invisible and frictionless.
The remaining twenty per cent (the heavy users) get free model choice, with their own framework and more personal responsibility. The budgets are profile-based, not uniform: linked to the value that AI adds to a role. An analyst who generates financial models daily justifies a higher budget than someone who uses AI sporadically.
This shifts you from controlling to managing. You don't impose a flat ceiling but invest consciously where the value lies.
Where the CFO manages for value, the CIO manages for the architecture that makes that possible. Model routing (the right task to the right model) only works if your architecture allows it. There's a pitfall here: a carefully configured combination of models and routing rules can be outdated within months. New models become available, with different prices, a different way of controlling and different trade-offs.
The core question for the CIO is therefore not 'which model is best now?', but: can we switch model or provider without rebuilding the entire solution? Do we have a multi-model strategy? Do we see token consumption per model, per work process and per business unit? If you design your architecture for freedom of choice instead of for the specifications of one provider, you can move with the market without getting stuck in costly dependencies.
The first wave of AI was human-driven: one prompt, one answer. Visible, manageable. Agentic AI changes that. Agents plan, reason, retrieve data, call tools and trigger work processes, with little human intervention.
One simple request – 'create a supplier risk report' – has one or more agents search purchasing data, retrieve contracts, analyse market signals and compile a report. Each step consumes tokens. AI costs are then no longer 'costs per prompt', but costs per autonomous task, per work process, per completed outcome. Consumption disappears and embeds itself in your processes.
That requires governance on work processes instead of on prompts: which tasks may an agent perform, which data access, how many calls make before a human intervenes, what is the maximum cost price per task? And it requires boundaries against so-called 'runaway agents', which keep searching without better results. Every AI system needs a licence to act, with explicit boundaries.
Token consumption is no longer a technical detail that belongs with IT. It's a governance issue that CFO and CIO tackle together: the CFO links consumption to value, the CIO builds the architecture that makes that possible.
The first step is small and concrete: have CFO and CIO jointly map out how many tokens are consumed per model, process and business unit, linked to the key use cases. That baseline measurement immediately shows where budget is wasted and where extra investment pays off.
But the core remains simple: it's not about how much you spend, but about what you get back per token. If you immediately reach for a ceiling, you choose the panic measure. If you choose transparency, profile-based budgets and a clear link with value, you transform AI from a rising cost item into a manageable lever.
The winners in 'enterprise AI' aren't those with the most powerful models. They're those who convert every token into measurable value. So don't start with the question 'how do we limit this?' Start with: do we know what each token delivers us?