Uber spent its entire 2026 budget for AI coding tools in the first four months of the year. Then it put every engineer on a $1,500 monthly allowance.

The numbers behind that are worth sitting with. Around 5,000 engineers, monthly usage rates between 84 and 95 percent, individual bills running anywhere from $150 to $2,000. The company’s CTO said in April they had already burned through the full year’s Claude Code budget. Adoption was not the problem. Adoption was total.

The problem was the other side of the ledger. Uber’s COO, Andrew Macdonald, went on a podcast and said the quiet part out loud: all that token spend does not yet map to shipped features. The productivity gain everyone assumed was there has not shown up in the numbers.

Read that as an auditor and something should click. Rapid, near-universal spend on something whose value nobody has measured, followed by a reactive cap once the bill landed. If this were travel and entertainment, or cloud, or consulting fees, your function would already have a workpaper open.

The pendulum swung in about eight weeks

Three months ago I wrote about the governance gap, and one of the things I flagged was token leaderboards. Meta was ranking employees by how many tokens they consumed, treating usage as a performance signal. The logic was that usage means experimentation and experimentation means improvement, so more tokens must be better. Uber ran the same play, with leaderboards encouraging people to compete for AI usage.

I asked at the time: what are we actually spending our tokens on?

The market has now answered, and the answer is “we’re not sure, so stop.” Meta is imposing centralized spending controls after internal consumption headed toward billions of dollars for the year. CNBC has a name for the reversal. We went from tokenmaxxing to efficiency, almost overnight. Others are calling the new mood tokenminimizing, or token discipline.

Part of this is timing. Both Anthropic and OpenAI filed confidentially for IPOs in early June, and their biggest customers are suddenly being asked by their own boards what they get for the spend. When you are preparing to be a public company, “we ranked people by how much they burned” is not a line you want in the S-1.

But the deeper point is that the whole thing whipsawed from one extreme to the other in a matter of weeks. First the instruction was use as much as you possibly can. Then it was here is a hard cap. Neither of those is governance. They are the two camps I described in April, the “only risk is not adopting fast enough” crowd and the “lock it down” crowd, and this time the same companies played both roles inside a single quarter.

Why this is our problem, not IT’s

It is tempting to file this under engineering budgets and move on. Do not.

What Uber and Meta are living through is a control failure of a very familiar shape. A new category of spend appeared. It grew fast. It had no meaningful approval gates, no consumption limits, no allocation to business units, and, critically, no way to connect the money going out to value coming back. Then it got capped in a panic when the invoice arrived.

Strip out the word “AI” and this is a textbook finding. Expenditure without a business case. Absence of budget controls. No link between cost and benefit. Reactive rather than designed limits. We have language for every one of those, and we have tested them for decades.

The reason AI spend escaped that scrutiny is that it was framed as strategic rather than operational. Nobody wanted to be the person slowing down the future. So the normal discipline that applies to any other line item got suspended, and it got suspended precisely in the area growing fastest. That is the pattern worth remembering. The risk was not the technology. The risk was that we exempted the technology from the controls we apply to everything else.

The measurement gap is the real story

Macdonald’s comment about spend not mapping to output is the sentence I keep coming back to, because it points at the thing almost no organization has built: a way to attribute AI cost to AI value.

This is squarely our territory. Connecting expenditure to outcome, testing whether a stated benefit actually materialized, and challenging the business case when it does not, is close to the definition of what internal audit does. The efficiency era is not bad news for the profession. It is the moment our core skill becomes relevant to the single fastest-growing cost in the enterprise.

And the timing lines up with the standards. COSO put out guidance this year on internal control over generative AI, and it reads the way you would expect: expect model documentation, validation evidence, change logs, records of human override. On the finance side, CFOs are now being asked to justify AI return on investment and attest to the reliability of AI-assisted reporting. The pieces of an assurance framework are arriving. What is missing is the practice of using them.

What I’d actually look at this quarter

I am wary of turning this into a checklist, because the real work is judgment. But if I were scoping this into an audit plan right now, I would start with a few concrete questions.

  • Can the organization produce total AI spend for the year, broken down by tool, team, and business purpose? If the answer takes more than a day, that gap is itself a finding.
  • Are there approval gates and consumption limits, or did controls only appear after a bill shocked someone? A cap imposed in month five tells you the design was absent in months one through four.
  • For any AI deployment with a business case, has anyone gone back to test whether the claimed benefit showed up? Not projected. Measured.
  • Where AI touches financial reporting, does the COSO-style evidence exist: model documentation, validation, change logs, human override records?

None of this is about spending less. Uber’s engineers using AI at a 95 percent clip is not the failure. Some of that spend is almost certainly creating real value. The failure is that nobody could tell which part.

That is the gap. And unlike a lot of what is happening in AI right now, it is one we already know how to close. This one is ours. We should take it before someone writes us a framework that assumes we can’t.


If you’re building AI spend controls or trying to connect that spend to value inside your function, I’d like to hear how you’re approaching it. Get in touch.

Sources and further reading: