AI Cost Overrun Audit: Token Governance & Spend Controls for Small Business
Why run an AI cost overrun audit?
60% of organizations using AI will face cost overruns related to the technology, caused by lack of usage tracking — and most companies (56%) are still implementing AI tools without clear usage policies, according to an April survey of 1,300 senior marketers conducted by Gartner (Digiday, Sept 4 2026). AI agents make the problem worse because they spend autonomously: they can run unattended, pick the most expensive model, and drift outside the brief they were given. An AI cost overrun audit puts per-agent budgets, spend alerts, audit logs, and human approval gates in place before the overrun shows up on an invoice.
Media agencies that deploy AI to speed up planning and buying are learning the same lesson every small business will: the technology won't deliver savings if it's left to run by itself — it has to be monitored. As agentic tools proliferate, agencies are building tracking and auditing tools so staff can check that AI tools aren't hallucinating or burning through tokens. "It can get out of control very, very quickly," Jonathan Whiteside, global EVP of technology at Dept, told Digiday.
Why cost overruns happen: the usage-tracking gap
Gartner's estimate is blunt: 60% of organizations using AI will face AI-related cost overruns from lack of usage tracking. The survey also found 56% of companies are implementing AI tools without clear usage policies, and marketing leaders in particular are less likely to assign financial controls to their team's AI usage. What you cannot see, you cannot control — and with AI, usage is invisible by default because it is measured in tokens, not seats.
Cost overruns are also a model-selection problem. More powerful models cost more per token, and "a lot of the [rising] token consumption costs are because people are literally not choosing the right model to meet the need of the activity," Gartner analyst Nicole Greene told Digiday. Defaulting to the most capable model for every task inflates spend even when nothing else changes.
What agencies already do: real-world cases
Several agencies are ahead of most small businesses on AI spend governance, and their controls map directly onto what an SMB can adopt. Performance media agency Rise, part of the Quad agency group, has been testing AI media-buying agents with a supermarket client since June, according to group director Klaudia Smykowska. Rise uses an audit-log feature developed by PubMatic — an SSP that has worked with agencies testing AI media-buying tools, including Butler/Till and Abovo Maxlead in the Netherlands — to monitor when agents are "drifting" outside pre-set parameters and to log the actions taken. "I can go and ask, 'why did you make that change? What was the thought process based on that initial brief?'... I can make sure everything is recorded and that we can reconstruct what happened and why," Smykowska said.
PubMatic's tool shows what a real audit trail looks like: "Any change [the agent] makes in any environment is time-stamped, stored with all the details of the change, and how that change was made — was it in the UI [for example] — is also recorded," said Harry Tong, PubMatic's director of sales engineering. Other agencies track differently: Brainlabs grants staff AI tokens on a tiered system and reviews requests for more, while Dept routes every prompt and agent through an internal "AI gateway" that picks the model centrally by commercial and legal criteria — after people "burned through 1.5 million tokens in a day," per Whiteside. PMG's "Alli For You" tool caps daily token usage per user and escalates anyone who routinely hits the cap to a human review. The throughline: agents and employees need bounded, visible, reviewable spending — as Whiteside put it, "every deliverable has an accountable human."
The AI cost overrun audit checklist for small business
Work through the seven items below. Each one has a concrete verification step. If an item fails, the fix is a setting, a written policy, or a vendor question — not a six-figure project.
-
Set per-agent token budgets. Give each agent, team, and (if they have individual logins) each employee a defined token or dollar budget per period, the way Brainlabs issues tiered token access and PMG caps daily usage. Decide the cap from the task, not from last month's bill. An agent that cannot prove its value within budget should be re-scoped, not re-funded automatically.
-
Turn on spend alerts at a threshold you will actually notice. Configure alerts on model usage, API spend, and per-agent consumption so you are notified at a warning level, not only at the hard cap. Rise's team monitors licensing and token spend continuously rather than waiting for an invoice. Alerting is cheap; a surprise invoice is not.
-
Keep an audit log of agent actions and changes. Require every agent tool you deploy to record what it changed, when, and how — the PubMatic audit-log standard: time-stamped, detailed, reconstructable. You should be able to ask an agent "why did you make that change?" and get an answer grounded in a log, the way Rise does with its media-buying agents.
-
Watch for drift outside pre-set parameters. Define the parameters an agent must stay within — budget, audience, channels, price ceilings — and review agent actions against them on a schedule. Drift detection is what turns a log from a forensic tool into a cost control: you catch an agent that wandered outside its brief while the overrun is still small.
-
Require human approval gates for increases and high-cost actions. Nobody should be able to raise their own token budget or approve their own agent's expensive action. PMG's cap review, Brainlabs' tier review, and Dept's central gateway all put a human between the request and the spend. The gate should ask whether the usage was useful — "we want people to spend, we just want them to do it usefully," said Brainlabs founder and CEO Daniel Gilbert.
-
Write and enforce a clear usage policy. The 56% stat is your warning: most companies run AI with no usage policy at all. Write down who may use which tools, which tasks may run on which models, what must have human approval, and who owns the budget. The policy does not need to be long; it needs to be real and reviewed at the same cadence as your other business policies.
-
Demand vendor pass-through price transparency. Ask every AI vendor and agency partner in writing: are you charging me the provider's listed model price, or a markup? Which models and token counts do I actually consume, and can I see itemized usage? When a vendor routes your work through its own model gateway, the model choice and the pricing should both be visible to you — the same transparency Dept gives its clients when it decides centrally which model a request uses.
The model-selection trap. The single cheapest fix in this checklist is also the easiest to skip: stop defaulting to the most powerful model for every task. Token costs scale with model capability, and agencies report that much of the inflation comes from using frontier models where a cheaper one would do. Assign model tiers to task types in your usage policy, and let the approval gate question any request for the most expensive tier.
Where this audit fits with your other AI checks
Cost overruns are a financial-governance problem, distinct from the security and code-provenance risks covered by the AI agent risk checklist and the code supply-chain controls on this site. If your agents can spend money on outside services at all, pair this audit with the agent-spend governance audit, which covers approved merchants and spending caps for agent purchases. And if your AI usage is paid through cloud or API keys, keep the key-hygiene controls in LLMjacking: How Leaked AWS Keys Burn Your AI Budget in place so a leaked key cannot quietly inflate the same bill you are trying to control.
Bottom line
The Gartner number is the business case: 60% of organizations using AI will face cost overruns from lack of usage tracking, and most companies have not even written a usage policy. You can run the seven-item AI cost overrun audit above in an afternoon: set per-agent budgets, turn on alerts, keep audit logs, watch for drift, require human gates, write the policy, and demand vendor pricing transparency. Agencies are building bespoke tools for this because the stakes are real — but for a small business, the controls are mostly settings, policies, and questions you already have the right to ask.
Frequently asked questions
What is an AI cost overrun audit?
An AI cost overrun audit reviews how your business tracks and controls what it spends on AI models, APIs, and agents: who and what consumes tokens, which tasks run on expensive models, whether agents can act outside preset budgets, and whether vendors pass through model costs transparently. The goal is to catch overruns before the invoice arrives.
What percentage of companies face AI cost overruns?
Gartner estimates that 60% of organizations using AI will face cost overruns related to the technology caused by lack of usage tracking. The same April survey of 1,300 senior marketers found most companies (56%) are still implementing AI tools without clear usage policies.
What is an AI token budget?
An AI token budget is a spending limit on the tokens a person, team, or agent may consume in a period. Agencies testing AI buying agents give each agent its own budget or cap, and a human approves increases after reviewing whether the usage was justified.
How do AI agents cause cost overruns?
AI agents cause cost overruns by running unattended, choosing more powerful and expensive models than a task needs, drifting outside the parameters of their original brief, and burning tokens on loops and retries. Department heads interviewed by Digiday described people burning through 1.5 million tokens in a day and said agent costs can get out of control very, very quickly.
What is AI agent drift detection?
AI agent drift detection is a control that watches whether an agent is acting outside the pre-set parameters of its brief. PubMatic's audit-log feature, tested by agencies including Rise, Butler/Till and Abovo Maxlead, records every change an agent makes with a timestamp so a human can ask why a change happened and reconstruct the decision.
What should an AI cost overrun audit checklist include?
A practical AI cost overrun audit checklist covers: per-agent token budgets, spend alerts, audit logs of agent actions, drift detection outside pre-set parameters, human approval gates for increases, clear written usage policies, and vendor pass-through price transparency for the models you actually consume.
Source: Sam Bradley, Digiday — "Media agencies build audit tools to prevent AI agents from overcharging" (September 4, 2026, digiday.com). Gartner figures and agency case details as reported by Digiday from Gartner's April survey of 1,300 senior marketers.