AI Cost Optimization
AI bills are racking up and people are awful at spending money wisely
Hopefully by now, everyone should have realized that Anthropic and OpenAI evangelizing developers to run 19 hour-long coding jobs, without putting any thought into hard ROIs, was an elaborate psyop to juice revenues before their IPOs.
Jokes aside, AI cost management — and more broadly, financial planning around token spend — are fast becoming massive problem areas for enterprises. At least, Zuck feels that pain, and Meta is reportedly building internal tools to track AI token spends firm-wide to better measure ROI on AI spend. Meanwhile, Uber is introducing per-employee budget of $1,500 a month, indicating that AI spend policies will have hierarchical limits (per business unit, per employee).
What’s funny is, these moves from Uber and Meta remind me, a former trader, of how risk management is done in Wall St banks and hedge funds. In finance, the “ingredient” to make money is essentially just… money. So it’s important to watch it like a hawk to prevent traders from blowing up your firm. Similarly, as AI takes increasing share of firm’s final output, tech industry will soon learn to do the same for AI tokens.
For example, if you are a stock trader at a hedge fund, you have $X million that you can lose each week, month, quarter, etc. If you lose more than that, they keep you on a tighter leash. If you blow through your risk limits too often without showing any results (making money, which is foreign to many tech workers), you are let go. If you try to mess with your risk limits by committing fraud, you are fired on the spot.
And in some sense, managing tech bros using AI agents for coding has deep parallels with managing traders, and the whole field sorely needs some concept of “risk limits” to constrain their behavior regarding AI usage.
Here’s how to think about this. Essentially, every developer using AI is suddenly a capital allocator (or a trader), since tokens are capital. More broadly, as “AI tokens” take up a greater fraction of a firm’s end output (versus labor), then every worker becomes a capital allocator.
So, if you are a business, you are effectively hiring developers as discretionary traders to make money with the capital stock that you are providing. Without risk limits or policies or compliance, you get an unconstrained shoveling of money into OpenAI or Anthropic’s coffers, straight from your wallet.
Unfortunately, companies are still mistakenly seeing AI agents as yet another IT spend. It’s NOT. It should be viewed as a capital allocation decision, especially vis a vis outsourcing. Arguably, you can’t put a clean, fixed budget on a spend that can increase geometrically and unpredictably. Instead, things should be aligned with work.
If the Wall St analogy holds, every enterprise’s FinOps need to be upgraded with things like:
fat finger protection: prevent a rogue employee from running 50 hour background jobs with the latest fancy token sucking feature in Codex
P&L measurement: measuring how much value was delivered for an agentic workflow that costs $X to run.
compliance and fraud protection: as tokens get tied directly to real dollars, you need guardrails against abuse — shadow AI, shadow API keys, employees burning company tokens on personal side projects, or data quietly leaking out through prompts.
None of this is exotic. It’s the same middle-office plumbing every trading floor already runs, and the moment enterprises decide they need it, they will adopt it. We’ve seen this movie before. It was called cloud Fin Ops.
Note, if you are an enterprise that’s spending $10M+ on tokens and need cost optimization help, DM me as I have developed a framework to slash costs dramatically. Or stay subscribed for more info in the coming weeks.
Anyways, between 2018 and 2023, a hot topic in enterprise cloud was “cloud cost optimization / cloud FinOps”, and a slew of companies were founded to help enterprises lower their AWS / GCP / Azure bills in exchange for some cut. Almost none of them exited for more than $1Bn, but still many decent businesses were founded helping companies save money on cloud.
For example, say you are a Fortune 100 bank, and you are spending $80M a year on S3 (AWS object storage) storing petabytes of data that’s growing 30% a year. To mitigate this problem, you hire a company to scan every S3 bucket, and detect / fix any wasteful configurations, such as storing archive data in expensive storage tiers. The cost savings were immediate and measurable, which is why “cloud cost optimization” startups did well for a while, until their feature sets got absorbed (read: copied) into AWS’s roadmap.
Fast forward to 2026, I think there’s a pretty wide-open market for providing products and services around “AI cost optimization” as well as “AI FinOps”, given the recent hoopla about exploding AI costs. Which is probably a premature worry, honestly, given that AI tokens represent less than 2% of total expense even at the most “AI-native” companies.
But it’s inevitable that CFOs start demanding tighter governance and monitoring around AI-token costs, and feed that data back into financial planning, especially as AI adoption increases over time.
This necessitates better tools to monitor and optimize AI token bills, as well as generate real time ROI calculations to assess the value of every token used inside the company.
For example, right now, no company is currently able to answer basic questions such as:
which business unit is delivering the most “value”, adjusted for AI token spend?
which models account for the most $ nominal token spend? Are we inadvertently using Opus 4.8 to do something basic like writing unit tests?
how do we prevent runaway spend *before* the bill lands, instead of explaining it after the fact?
So what does the AI FinOps stack actually end up looking like? My bet is it splits the same way Wall Street and cloud both did.
On one side, the “risk and P&L” layer — the unglamorous, CFO-facing system of record that maps every token to a business unit, an owner, and a workflow, then ties it back to the value it produced, so finance can decide who gets more rope and who gets less.
On the other side, the execution layer — smart compaction, prompt compression, model routing and hot-swapping, ideally running right at the inference layer. Think of it as smart order routing in stock exchanges but for AI: genuinely valuable, but the kind of thing that gets commoditized or absorbed into the platform, the way AWS quietly ate the cloud-optimization startups.
But my advice to enterprises is, stop worrying about trying to do everything at once. Instead focus your efforts on enforcing basic hard limits per employee, taking an inventory of all API keys issued for every business unit, and visualizing how much every business unit is spending in aggregate on a trailing 30 day basis.
Everyone is interpreting Meta, Amazon, and Microsoft’s pushback on AI token spending as an evidence of AI not delivering value. I think that’s the wrong framing. AI and capital are mostly amplifiers. If you give money to a bad trader, he or she will lose money, and that’s the same with AI tokens.
The winners won’t be the companies that spent the least; they’ll be the ones that treated tokens like capital, built themselves a real risk desk for AI.
About Me
I write the “Enterprise AI Trends” newsletter (read by over 60K readers worldwide per month). Previously, I was a Generative AI architect at AWS, an early PM at Alexa, and the Head of Volatility Index Trading at Morgan Stanley. I studied CS and Math at Stanford (BS, MS).



Great work as always John. I worked for a Cloud FinOps company in 2018, I can see a need for an AI FinOps solution.
Love this: “Warning: I may not respond if world cup matches are too good.”
Some good matches already and more coming.