Skip to content
[ blog post ]

AI token cost: why the bill outruns the budget

Atmospheric deeplit® hero in Midnight Blue and Neon Purple with the build-vs-buy icon and the title 'AI token cost: why the bill outruns the budget'.

July 4, 2026

H. Kamkar

AI token cost is falling. The bills are exploding anyway.

deeplit® builds private AI infrastructure for Dutch technical founders and CTOs selling AI features into regulated industries: hospitals, banks, insurers, government agencies, law firms, pharma R&D groups. There are two ways to run it. On the customer's own hardware, as deeplit® Private. Or on single-tenant GPUs that deeplit® provides, billed at a flat hourly rate, never per token, as deeplit® Cloud. This piece is about the second thing, cost, because in 2026 AI token cost became the fastest way for a company to lose control of its own budget.

Start with the paradox, because it is the whole story. The price of a million tokens has fallen by more than 99 percent in three years on the cheaper model tiers. And yet the FinOps Foundation's State of FinOps 2026 report, drawn from more than a thousand organizations, found that 73 percent of enterprises blew past their AI cost projections. Companies were reporting, in April, that they were already three times over their entire 2026 token budget. The unit got cheaper. The bill got bigger. Both are true at once, and the gap between them is where founders are getting hurt.

Two numbers frame the range. A fintech told its staff to lean harder on AI coding, and one employee spent $81,267 in credits in about a week building a meme game. At the other end, an AI consultant told Axios that a client ran up around $500 million on a token-billed tool in a single month after leaving usage uncapped. One person on a joke, five figures. One company on autopilot, nine. Neither number was a pricing error. Both were the pricing model working exactly as designed: when the meter runs on tokens, cost is a function of how much the software is used, not how much value it returns.

That is the thing to understand about AI token cost. It is not a budgeting problem you can dashboard your way out of. It is a pricing-model problem, and pricing models can be changed. Cost is the lead argument for private AI, ahead of privacy, because it is the one every founder feels in the same month it happens.

When the meter runs on tokens, cost is a function of how much the software is used, not how much value it returns.

How AI became a metered utility

What a token is, and why you are billed for it

A token is the unit a language model reads and writes in. Roughly four characters of English, about three-quarters of a word. Every prompt you send is counted in tokens, every word the model generates is counted in tokens, and the providers bill you for both directions. It is a clean unit for a vendor to meter, the way a utility meters kilowatt-hours. The catch is that, unlike electricity, you do not control how many tokens a task consumes. The model does, and increasingly another model does it on your behalf, thousands of times a minute.

For most of the generative-AI era that did not matter, because you were not really paying for it.

The subsidy that hid the real cost

For about three years, AI tools were priced like a gym membership: a flat monthly fee for effectively unlimited access, underwritten by venture capital that wanted market share more than margin. That subsidy hid the real cost of inference. Some heavy users generated more than $35,000 of compute against a $200 monthly subscription, a subsidy of more than a hundred times, absorbed by providers racing to lock in customers. The companies selling inference have been pouring capital into data centers at a rate that is close to doubling year over year, building for a demand curve they can see coming but cannot yet profitably serve.

That could not hold, and in early 2026 it broke. The large model providers moved their enterprise customers onto token-based billing. One widely used coding tool flipped every plan to usage-based credits on a single day in June, and heavy users reported bills 25 times higher than the month before. The flat-rate era did not end because anyone chose it. It ended because the subsidy ran out, and the true cost of metered inference landed on the customer all at once. Diffuse, hidden budget lines became measurable, per-task costs overnight, and the per-task costs were larger than anyone had modeled.

Why AI token cost outruns every budget

Total spend is price times volume, and volume is winning

Hold the arithmetic in your head: total spend equals the cost per token times the number of tokens. Every headline about AI cost per token falling is describing the first term. Every runaway bill is the second term winning anyway. The cheapest tiers will keep getting cheaper; that is not in question. In the same breath, Goldman Sachs projects that agentic AI could multiply token consumption twenty-four fold by 2030. When volume grows faster than price falls, the bill goes up while the sticker price goes down. That is not a paradox once you write it as multiplication. It is the default.

A better rate on cheaper LLM pricing, applied to a workload quietly consuming ten times more tokens each quarter, is a smaller discount on a bigger number. The variable that actually moves your bill is the one the vendor does not control and, on a metered plan, neither do you: how many tokens your workloads burn.

Agentic workflows are token amplifiers

The thing driving volume is the shift from chat to agents. A single-shot question sends one prompt and gets one answer. An agent plans, calls a tool, reads the result, replans, calls another tool, and loops, sometimes for dozens of steps, before it returns anything to a human. Every step is billed.

The multipliers are not small. One 2026 analysis of agentic coding found that a single agent task can consume on the order of thousands of times the tokens of a one-shot question. Orchestration overhead alone runs one to three thousand tokens per workflow before the model does any real work, and a 200-token user request can expand past 40,000 tokens once retrieval and tool chatter are counted. Reasoning models make it worse, generating intermediate deliberation tokens at every step. That is what agentic AI cost looks like under the hood: retry loops, growing context windows, background calls, and multi-agent chains, none of it visible until the invoice arrives. The workload that demos beautifully in a sprint is the same workload that arrives as a five-figure line item at the end of the quarter. deeplit® wrote separately about running agentic workflows on internal infrastructure; the cost behavior is a large part of why the topic keeps coming up.

The incidents are the model working as designed

The blow-ups

The 2026 examples are not cautionary outliers. They are what the meter does at scale, to sophisticated teams.

Uber burned through its entire 2026 AI budget in four months. Adoption of a token-billed coding assistant went from a third of its engineers to 84 percent in a single month, the tool worked, and the company capped each engineer at $1,500 a month per coding tool because the run rate was otherwise unsurvivable. Its chief technology officer described being back at the drawing board. This is a company with a world-class finance function, and it could not forecast the bill one quarter out.

Microsoft handed thousands of developers a token-billed coding assistant, watched the usage, and pulled most of the licenses about six months later, moving one division onto a fixed-seat alternative by a June 30, 2026 deadline. The per-engineer bill had run as high as roughly $2,000 a month. A company that sells AI curbed its own internal AI use on cost grounds.

Slash, a fintech, told its staff to lean harder on AI coding. One employee spent $81,267 in credits in roughly a week on a meme game. At a travel company, a routine renewal of one coding tool came back four to five times more expensive than the year before. And the consultant's client that reportedly spent around $500 million in a single month, unnamed and unconfirmed, was not dismissed as impossible by anyone in the industry, which tells you where the ceiling now sits.

The common thread is not incompetence. It is that on a per-token plan, cost is decoupled from value. A person can burn five figures on something that produces nothing. An organization can burn nine before finance closes the month. The meter does not know the difference, and neither does your budget until the invoice arrives.

When the agent costs more than the person it replaced

Here is the line that should stop a founder cold. Bryan Catanzaro, Nvidia's vice president of applied deep learning, told Axios that for his team, the cost of compute is far beyond the costs of the employees. Nvidia sells the compute. Gartner forecast in June 2026 that AI coding costs will pass the average developer's salary by 2028. A quarter of technology leaders already spend between $200 and $500 per developer per month on coding tokens, and about 6 percent spend more than $2,000. Fortune, Axios, and Forbes all ran versions of the same headline through the spring: the tool now costs more than the person it was supposed to replace.

The joke that a human is cheaper than an agent is half right, and its conclusion is wrong. The agent does not cost more than the human because AI is inherently overpriced. It costs more because you are renting it by the token, so the cost scales with how hard you use it instead of with the value it produces. Use it lightly and the economics look fine. Use it the way you actually want to, everywhere, all day, and the meter turns your best case into your worst bill. Change how you pay for it, and the comparison changes with it.

Why the standard fixes only cap the symptom

The industry's response has been to instrument the meter. There is now a whole discipline of token-level financial operations. Spend-management platforms have moved into AI cost tracking, observability vendors have added token monitoring and GPU dashboards, and a wave of model routers now send cheap tasks to small models and reserve the expensive models for hard ones. The Linux Foundation has announced plans for a Tokenomics Foundation to standardize the practice. These tools are real, and they help. Gartner notes that intelligent routing alone can cut a coding bill by 40 to 85 percent.

But look at what they actually do. Routing, caps, tiering, and observability lower the number, and more to the point, they leave it variable. You still cannot tell finance what next quarter costs, because it still depends on how much your teams use the thing you bought so they would use it. A cap is not a forecast. It is rationing the tool you adopted for leverage, which means throttling the productivity you were paying for the moment it gets valuable. And the observability layer is itself a recurring cost, a tax you pay to watch a meter you do not control. Optimization makes a runaway bill slightly less runaway. It cannot make it predictable, because you cannot forecast what you do not own.

Blueprint of a coastal lighthouse with its lantern room glowing purple and a warm coral point on a distant ship, captioned 'The lamp burns the same for one ship or a thousand.'

Flat rate, never per token: what owning the GPU hour changes

deeplit® was built around the other side of that sentence. Instead of paying for inference by the token from a shared pool, you run open-weight models on capacity that is yours. Two shapes. deeplit® Private puts the deployment on the customer's own server or data center, setup plus a monthly platform license. deeplit® Cloud puts it on single-tenant GPUs that deeplit® provides, billed at a flat hourly rate, never per token, with a same-day start. Same platform, same models, same code. What changes is the meter: you pay for a machine by the hour, not for tokens by the million. A heavy day and a light day cost the same. Finance gets one number it can forecast before you run anything.

Why a flat rate is honest here, not a loss-leader

It is fair to be suspicious of a flat rate, because you just watched one collapse. The distinction is what sits underneath it. The gym-membership subsidy lost money because shared inference has no ceiling: the heaviest user consumes without limit and the provider eats the difference. A dedicated GPU is not like that. It does a bounded amount of work per hour whether you push it hard or leave it idle. Pricing a bounded resource at a flat rate is not a bet that you will underuse it. It is renting the machine for the hour, the way you would rent any other capital equipment. That is why the flat rate is durable here and was not durable there. It prices the hardware, not how hard you use it.

What the flat rate gives up: you provision capacity

This is the honest cost of the model, and it matters. A flat GPU-hour rate buys a bounded amount of capacity. A given GPU serves a finite number of requests per second. If your load spikes past what you have provisioned, you do not get a bigger bill, you hit a ceiling, and scaling means adding GPU capacity you plan for in advance. The token economy makes that provisioning problem disappear by pooling it: the provider holds spare capacity across thousands of tenants and rents you a slice on demand, so a sudden spike is absorbed by the pool instead of by your planning. That elasticity is a real advantage of shared metered inference, and it is exactly what you give up when you take dedicated capacity. You trade the ability to absorb any spike instantly for the ability to know your bill in advance.

What this looks like in real numbers

Take a workload we ran with a data-analytics team, under an NDA, so no names. The job was to analyze a large volume of public business data: scraped company websites, professional profiles, chamber-of-commerce records, piles of unstructured text. Batch analysis that is pure token consumption from the first request to the last.

It ran an open-weight model on a dedicated GPU, flat out for 48 hours, across more than two million requests. Each request filled an 8,000-token window, roughly 60 percent input (the business data fed in) and 40 percent output (the analysis written back). That is on the order of sixteen billion tokens: about 9.6 billion in and 6.4 billion out. The GPU cost four dollars an hour, so the whole job came to $192, and they knew that number before it started. They bought tokens at just over a cent per million.

Now rent the same work by the token, where input and output are billed separately and output costs several times more. Three points of comparison on the same 9.6 billion tokens in and 6.4 billion out, at current rates:

  • The same open-weight model, hosted per-token (about $0.20 per million in, $0.60 out): roughly $5,800. This is the honest comparison, the identical model on both sides, and it is already about thirty times the flat cost.
  • A mid-tier hosted model (about $1 in, $3 out): roughly $29,000, around a hundred and fifty times the flat cost.
  • A frontier model (about $5 in, $10 to $30 out): between $112,000 and $240,000, most of it output.
Bar chart: the same 16-billion-token job costs $192 on a flat GPU-hour versus about $5,800 rented per token, roughly thirty times more.

The work does not change. Only the meter does.

The shape of the two bills matters more than the multiple. The $192 was fixed and known in advance; run the job again next week and it is $192 again. The per-token figure you learn at the end, and it grows every time the dataset does. This was a steady batch workload with a knowable peak, which is exactly where owning the hour wins.

Match the pricing model to the shape of the workload

So the decision is not flat-rate-good, per-token-bad. It is a question about the shape of your load. Flat-rate ownership wins when your usage is steady and forecastable enough that provisioning for your peak is cheaper than renting elastic burst you rarely touch. That describes most of the regulated internal AI deployments deeplit® sees: retrieval-augmented generation over internal documents, document extraction and classification at volume, internal copilots, agentic workflows over internal tools. Predictable, all-day, load-bearing work whose peak is knowable. The token economy wins for the opposite shape: spiky, bursty, consumer-scale traffic where demand swings by orders of magnitude and you would leave most of your provisioned capacity idle. deeplit® says this plainly, because a vendor that pretends its model fits every case is not worth trusting. If your workload is genuinely spiky or consumer-scale, or you need the newest closed frontier model, a per-token public API is the honest answer. If it is steady, you are paying a variable premium for elasticity you are not using.

I have watched the wrong side of this from the inside. A few years ago I ran a small financial-data startup on a per-token chat API with retrieval over our own documents. The product was good. The bill was not something we could plan around. It moved with every customer we onboarded and every feature we shipped, and the better the product did, the more the economics fought us. Owning the hour would have turned a number we feared into a number we set. That experience is a large part of why deeplit® exists, and it is the same conversation I now have with founders pricing the real cost of self-hosting an open-weight LLM against a metered API, or weighing the private LLM vs public LLM decision one level up.

If your token bill has stopped being forecastable, that is the signal to run the numbers on a flat rate. Book a deployment call. Bring your current usage and your worst month, and we will map it against dedicated capacity honestly. If per-token is the right answer for your workload, we will tell you.

If your finance lead is trying to forecast an AI line that moves every month, this post is written to be forwarded to them. The pricing-model argument is theirs to use in the budget conversation.

Token prices are falling. The bills keep exploding, because volume is winning.

A flat GPU-hour is a number you set before you run anything. A per-token bill is a number you learn at the end.

Frequently asked questions

Why do AI bills go over budget even as token prices fall?

Because total spend is price times volume, and volume is growing faster than price is falling. Cheaper per-token rates are being swamped by agentic workloads and heavier adoption pushing token consumption up far faster. A lower cost per token on a much larger number of tokens is still a larger bill, which is why 73 percent of enterprises exceeded their AI budgets in the FinOps Foundation's 2026 report.

What is the difference between per-token pricing and flat-rate GPU-hour pricing?

Per-token pricing meters every word your models read and write on shared, multi-tenant infrastructure, so your cost scales with usage and is hard to predict. Flat-rate GPU-hour pricing rents you dedicated capacity by the hour: a heavy day and a light day cost the same, and you know the number before you run anything. One is elastic but variable, the other bounded but forecastable. On cost, the more you use a dedicated GPU the lower your effective cost per request, because the hourly rate is fixed while the work it does rises; per-token runs the other way, with no ceiling. For steady workloads at real utilization the flat rate usually costs less and is always more forecastable. For light or highly intermittent use, per-token can be cheaper, which is why the honest answer depends on the shape of your workload.

Why are agentic workloads so much more expensive per task?

Because an agent does not send one prompt. It plans, calls tools, reads results, and loops, and every step is billed. Orchestration overhead alone can add thousands of tokens per workflow, and one 2026 analysis found a single agentic coding task can consume thousands of times the tokens of a one-shot question. Reasoning models add further intermediate tokens at each step. Agentic AI cost is driven by this multiplication, which is why agentic pilots so often produce the surprise bills.

Isn't a flat rate just another subsidy that will reprice later?

Not when it prices a bounded resource. The flat rates that collapsed in 2026 were subsidies on shared inference with no ceiling, where the heaviest users consumed far more than they paid for. A dedicated GPU does a bounded amount of work per hour regardless of how hard you push it, so a flat hourly rate on it is simply the rental price of the machine, not a bet that you will underuse it. That is the structural reason it holds.

What happens when my usage spikes past the GPU capacity I am paying for?

You reach the ceiling of that capacity rather than an unbounded bill, and you scale by provisioning more GPU capacity. This is the genuine trade-off of dedicated infrastructure: you plan for your peak instead of having a shared pool absorb it automatically. For steady workloads with a knowable peak this is straightforward. For wildly spiky or consumer-scale traffic, the elasticity of a per-token public API is the better fit.

When is a per-token public API still the right choice?

When your traffic is spiky, bursty, or consumer-scale, when your usage is light or occasional, or when you genuinely need the newest closed frontier model. deeplit® is built for steady, predictable workloads where open-weight quality is enough and a forecastable bill matters. If that is not your shape, a metered API is the honest answer, and a good deployment conversation should tell you so.

Read next

H. Kamkar headshot

H. Kamkar

Building private AI infrastructure for startups selling into regulated industries.