Physical Address
304 North Cardinal St.
Dorchester Center, MA 02124
Physical Address
304 North Cardinal St.
Dorchester Center, MA 02124

In April, Uber's chief technology officer told The Information that the company had spent its entire 2026 artificial intelligence budget. It was four months into the year. Nobody had misused anything. Engineers had used Claude Code for precisely the work it was built to do: parallel agent execution, large scale refactoring, automated test generation. Praveen Neppalli Naga reported burning $1,200 in a single two hour session during a personal demo of the tool.
In April, Uber’s chief technology officer told The Information that the company had spent its entire 2026 artificial intelligence budget. It was four months into the year. Nobody had misused anything. Engineers had used Claude Code for precisely the work it was built to do: parallel agent execution, large scale refactoring, automated test generation. Praveen Neppalli Naga reported burning $1,200 in a single two hour session during a personal demo of the tool.
By June the company had imposed a $1,500 monthly cap per employee per agentic coding tool, with an internal dashboard so staff could watch themselves approach the ceiling.
This is worth sitting with, because it contradicts something the market believes. The consensus view is that inference costs are collapsing, that Chinese open weight models are commoditising the tier below the frontier, and that cheap intelligence is arriving on schedule. That view is correct. It is also entirely compatible with enterprises detonating their annual AI budgets in a third of a year. Understanding why those two facts coexist is the most useful thing an investor can do with this cycle right now, because the reconciliation tells you which part of the AI trade is actually load bearing and which part was never more than a story about labour costs.
Start with the number that does not fit the deflation narrative. Silicon Data publishes an LLM Token Expenditure Index, a blended measure of what the market actually pays per million tokens rather than what any single model lists. J.P. Morgan Private Bank charted it against Bloomberg data and found that the index began a little above $1.00 in December 2025 and peaked just above $2.00 by mid June 2026. Not a straight line. It climbed to roughly $1.75 by early February, fell back towards $1.25 in March, then rebounded through April and May.
The effective price of intelligence, as bought by real buyers, roughly doubled in six months.
The resolution is that there is no longer one market. There are two, and they are moving in opposite directions.
BenchLM tracks live pricing across 143 models. Over the twelve months to August 2026, its index shows frontier token prices up 36.4 percent year over year while mid tier prices fell 35.8 percent. Open weight models now carry a median blended price 82 percent below proprietary equivalents, roughly $0.53 against $3.00 per million tokens at a conventional input to output ratio. Silicon Data’s segment readings tell the same story from a different angle: frontier inference at $4.20 per million tokens, open weight at $0.85.
Both halves of the argument are true simultaneously. The cheapest API achieving GPT-4 class quality now costs a small fraction of what GPT-4 itself cost at launch, and comparable frontier intelligence is several times cheaper than it was in early 2023. Yet the top tier has been repricing upward through 2026 as each new generation commands a premium for expanded capability. Cheap is getting cheaper. Expensive is getting more expensive. The blended index rose because buyers kept choosing the expensive half.
Price is only half of any bill. The other half is volume, and volume is where the real change happened.
The FinOps Foundation identifies three mechanisms, all structural rather than behavioural. Reasoning models bill for internal thinking tokens that never appear in the output. Context inflation means agentic workflows routinely push more than 50,000 input tokens per step through architectures originally designed for requests of a few hundred. And agentic loop multiplication means a single instruction can trigger twenty or more model calls before it returns anything to a human.
Put those together and per unit deflation becomes irrelevant. A workflow that costs 70 percent less per token but consumes 200 times more tokens per completed task is not a saving. Fortune framed the paradox precisely in its coverage of Uber: increasing AI use comes with higher costs even as per unit AI pricing falls.
Enterprises did not get their forecasts slightly wrong. They forecast the wrong variable.
Corporate culture then poured accelerant on the fire. In March, on the All-In podcast, Nvidia’s chief executive offered a thought experiment about an engineer earning $500,000 a year, and said he would be deeply alarmed if that engineer had not consumed at least $250,000 in tokens. Meta’s chief technology officer told a San Francisco summit his best engineer was spending his salary equivalent in tokens and called it easy money. At Sendbird, top consumers were ranked from Beginner up to a tier called AI God, defined as anyone burning at least 100 million tokens daily.
Tokenmaxxing was not a meme. It was a measured KPI. Uber ranked engineering teams on internal leaderboards by total tool usage, which is the proximate reason the budget vanished by April. At Meta, an employee built an internal dashboard nicknamed Claudeonomics that ranked more than 85,000 staff by consumption, with titles like Cache Wizard. In one thirty day window Meta employees collectively consumed roughly 60 trillion tokens. The leaderboard was shut down shortly after Fortune reported it.
There is an obvious conflict of interest in the loudest voice urging maximum consumption, given who sells the hardware the tokens run on. But the honest reading is that the status game was a symptom rather than a cause. Remove every leaderboard and the agentic loop multiplication problem remains exactly where it was.
Here is the part with money attached, and it is not the bubble question.
The entire labour displacement thesis rested on a cost claim: that a model doing a job costs a fraction of a person doing the same job. Goldman Sachs tested that claim across job types, and the results split hard. A coding agent came in at roughly $13.39 per day against something near $300 for a human developer. That gap is enormous and the substitution logic holds comfortably. But for a contact centre agent, Goldman put the daily AI cost at $92.90 against roughly $90 for the human.
That is not a saving. That is parity, before accounting for repeat contacts, escalations, or the interactions a model handles adequately but not well. Voice workloads are expensive because each turn requires transcription, retrieval, reasoning, generation and synthesis fast enough that the caller does not think the line has gone dead. Every stage bills.
So the substitution argument survives in code generation and dies somewhere around the middle of the job distribution. Which is a problem, because the headline forecasts about AI and employment were built on the assumption that it held everywhere.
Uber’s own president put the productivity side of it plainly, saying the link between rising Claude Code usage and consumer facing improvements is not there yet. That is a company with roughly 11 percent of live backend code changes written by agents, still unable to draw the line to shipped product.
Two claims circulating alongside this story do not survive checking, and MMI readers should be able to tell them apart.
The first is that data centre cancellations prove demand is fake. Bloomberg reporting, drawing on Sightline Climate, does put somewhere between 30 and 50 percent of US capacity planned for 2026 at risk of delay or cancellation. But the binding constraint is transformers, switchgear and batteries, not absent customers. Lead times for high power transformers have stretched from around 24 to 30 months before 2020 to as long as five years. SemiAnalysis has also argued convincingly that the headline figure is substantially a measurement artefact, because the projects being flagged sit overwhelmingly in a speculative pre construction bucket that was never going to land on a 2026 timeline. Physical bottlenecks are real. They are not evidence of demand collapse.
The second is that rising token costs mean companies will rehire the people they let go. Nothing in the data supports that yet. What the data supports is narrower and more interesting: that the cost case for substitution was overstated outside software engineering, and that finance departments are discovering this one budget line at a time.
Three things follow.
Revenue growth at the frontier labs is real but its quality is now a live question. Enterprises are not cutting AI spend. They are routing it, sending trivial calls to small models and reserving premium tiers for genuinely hard work. Practitioners report this cutting bills by well over half without measurable quality loss. Every dollar saved that way comes out of frontier revenue.
The governance layer is where the near term value accrues. Gateways, routing, caching, attribution and token budgets went from a niche concern to a board level one in about two quarters. Microsoft’s decision to move its Experiences and Devices engineers off Claude Code and onto GitHub Copilot CLI was officially framed as toolchain unification, though the cost dimension is not seriously disputed. When the company that owns a large stake in OpenAI and operates Azure is rationing agentic coding spend, the constraint is not confined to anyone’s edge case.
And the metered pricing shift changes what a software revenue line means. Per seat billing gave predictable revenue and predictable customer cost. Metered token billing gives neither. GitHub moved Copilot to credit based pricing in June. Uber caps at $1,500 per tool per employee. Walmart limited internal use of its own coding tool after adoption spiked. The industry is rebuilding cost control from scratch, in public, having shipped the products first.
This connects directly to the concentration and circular financing problems we examined in June, but it arrives from the opposite direction. That analysis looked at whether the suppliers could fund the buildout. This one looks at whether the buyers will keep paying premium rates once they learn what they are actually buying. The buildout thesis needs both answers to be yes.
The technology works. Engineers who use it do not want to give it up, which is precisely why the budgets blew. What has not been established is that it is cheap, and cheap was the entire argument.
This article is for informational purposes only and does not constitute investment advice. Conduct your own research before making investment decisions.