FinOps for Agents: Why AI Spend Behaves More Like a Trading Book Than a Software Licence
Agentic AI spend is variable, non-deterministic and has no natural ceiling, which makes it behave far more like a trading position than a technology purchase. This article looks at the evidence on what agents actually cost, the emerging discipline of FinOps for AI, and why the controls that fit are the ones financial services already knows how to run.
A cost that does not behave like a technology cost
Almost everything in a technology budget has a shape you can see in advance. A software licence is a number agreed for a year. A server has a capacity, and when it fills up someone notices. Even cloud, which broke the old capital expenditure model and needed a whole discipline invented to control it, is bounded by resources you can count.
Agentic AI is not like that. The same task, given to the same agent twice, can cost different amounts, because the agent decides how many steps to take. There is no instance to size and no licence to cap. The consumption is generated by the work itself, in real time, by a system exercising judgement about how much effort to expend. A misconfigured agent can spend a great deal of money in an afternoon, and the first anyone hears of it is the invoice.
That combination, variable consumption, real-time metering, and a limit that has to be imposed rather than discovered, is not new. It is the shape of a trading position. Financial services has spent decades building controls for exactly this problem, and almost none of that thinking has yet been pointed at AI spend.
What FinOps is, and why it exists
FinOps is the discipline that grew up around cloud for the same reason. The FinOps Foundation defines it as "an operational framework and cultural practice which maximizes the business value of technology, enables timely data-driven decision making, and creates financial accountability through collaboration between engineering, finance, and business teams".
The important word there is accountability. Cloud broke the traditional separation between the people who spend money and the people who answer for it: an engineer choosing an instance type is making a financial decision, usually without a finance conversation, and often without knowing the price. FinOps exists to close that gap, through visibility, attribution to a team, and continuous review rather than an annual budget cycle.
It is not a fringe idea. The Foundation was founded in February 2019 and merged into the Linux Foundation in June 2020. It describes a community of more than 120,000 people across more than 34,000 companies, including 97 of the Fortune 100.
The discipline has formally turned towards AI
In March 2026 the Foundation published an updated framework which, for the first time, treats AI as its own technology category alongside public cloud, SaaS, data centre and data cloud platforms. Its own description of what FinOps for AI addresses is a fair summary of the problem: "cost complexity, faster development cycle, spend unpredictability, and the need for a greater degree of policy and governance to support innovation".
The Foundation's own working group is more specific about why AI spend is harder to control than cloud spend. There are fewer established architecture patterns, so spending is more varied. Experimentation is higher, so any given service may be used briefly and abandoned. Purchasing channels are more diverse and less established. And, in a line that will be familiar to anyone who has watched a business function adopt a tool without telling anyone, the "low barrier to entry to become an 'AI Developer' means that people in non-technical roles will be operating as 'Engineers'".
On pricing, the Foundation is blunt: "Many AI models and services charge inconsistently, and may be purchased in many versions or variants. Pricing may change wildly up or down." That is a sentence with no equivalent in software licensing, and it is worth sitting with. The unit price of a critical input to your operations can be changed by a supplier, in either direction, without warning.
The Foundation's own numbers on adoption are striking. Its State of FinOps 2026 survey, with 1,192 respondents representing more than $83bn of annual cloud spend, found that 98% of respondents now manage AI spend, up from 63% in 2025 and 31% in 2024. AI cost management was ranked the single most needed skillset across organisations of every size.
Tokenomics, and the point at which a buzzword becomes an institution
The word that has attached itself to this is tokenomics, which is unfortunate, because it was previously a cryptocurrency term and carries some baggage. It is being used seriously nonetheless.
J.R. Storment, the Foundation's executive director, used it in his keynote at FinOps X in June 2026, describing tokenomics as "the emerging discipline of converting energy and capital into AI tokens, then consuming those tokens efficiently", and framing its purpose around two questions: "what does AI actually cost, and what is the value of intelligence?"
Earlier that month the Linux Foundation announced its intent to launch a Tokenomics Foundation, aimed at "establishing open industry standards, benchmarks, and best practices for the economics of AI infrastructure". Jim Zemlin, the Linux Foundation's chief executive, put the reason for it plainly: "Tokens have become the new unit of technology spend. Measuring and benchmarking token efficiency across different models and vendors is critical to how organizations make business decisions."
The list of supporting organisations includes Accenture, Booking.com, Flexera, Google Cloud, IBM, JPMorganChase, KPMG, Microsoft, Oracle, Salesforce, SAP and ServiceNow. A major bank putting its name to an open standards body for AI cost measurement is the clearest available signal that this is a live problem inside financial services and not a vendor invention.
What agents actually cost
The structural claim, that agents are expensive in a way chat is not, has a hard number behind it from a model provider's own telemetry. Anthropic's engineering team, writing about the multi-agent research system it built, reported: "In our data, agents typically use about 4x more tokens than chat interactions, and multi-agent systems use about 15x more tokens as chats."
The same piece contains the more interesting finding for anyone modelling this. Analysing performance on one benchmark, they found that "token usage by itself explains 80% of the variance", with the number of tool calls and the choice of model as the two other explanatory factors. Which is to say: how well the system performs is largely a function of how much you let it spend. That is a genuinely unusual property for a technology cost, and it is the reason a hard cap is a performance decision rather than only a budget one.
Their conclusion is the one to take into a business case: "For economic viability, multi-agent systems require tasks where the value of the task is high enough to pay for the increased performance."
The direction of travel is not towards this getting cheaper. In August 2026 Gartner predicted that AI inference costs per agentic workflow would increase more than fivefold through 2028, estimating that routing a task to an agentic reasoning model costs at least five times what the same task costs as a basic chatbot interaction. Will Sommer, a senior director analyst at Gartner, made the point that improving efficiency does not rescue the total: "Product leaders cannot rely on more efficient token economics to rationalize AI costs. Each successive generation of AI capability will necessitate more, and often more expensive, tokens." The failure mode he names is the one that matters here: "Defaulting to generic autonomous intelligence will result in unbounded costs orders of magnitude higher than those of optimized product ecosystems."
Unbounded is Gartner's word, not ours.
Deloitte reached the same place from the operational side in its Tech Trends 2026 report, published in December 2025: "As agents operate continuously, poorly configured agent interactions can trigger cascading actions like unpredictable resource consumption and ballooning costs, making cost management critical." Its prescription is explicit: "Organizations need specialized financial operations frameworks (or FinOps) to monitor and control agent-driven expenses and account for token-based pricing models."
The documented case
There is a great deal of anonymous folklore online about agents that ran in a loop and burned six figures overnight. Most of it traces back to content farms rather than to any incident anyone can name, and it should be treated accordingly.
The properly documented case is Uber. In May 2026 Fortune reported that the company had used up its entire 2026 budget for AI coding tools in four months. What makes it instructive is not the overrun itself but the reason Uber's chief operating officer, Andrew Macdonald, gave for finding the decision difficult. Speaking on the Rapid Response podcast, he described the problem of connecting the spend to anything: "it's very hard to draw a line between one of those stats and 'Okay now we're actually producing like 25% more useful consumer features'". His conclusion was that "if you're not actually able to draw a direct line to how [many] useful features and functionality you're shipping to your users, that trade becomes harder to justify".
The detail worth stealing is the mechanism. Uber had encouraged adoption using internal leaderboards ranking teams by usage. The incentive was to consume, and consumption is what it got. Anyone designing an AI adoption programme should look hard at what their own metrics are actually rewarding.
The metric that matters is not cost per token
The FinOps Foundation publishes the obvious formula, cost per token equals total cost divided by tokens used, and it is the least useful number in the discipline. It tells you what you spent, not whether it was worth spending.
The measure people are converging on is cost per outcome: the cost of one resolved case, one reviewed document, one completed reconciliation, one closed alert. The Foundation frames inference efficiency in exactly those terms, as "the cost to arrive at an anticipated outcome".
For a bank this framing has an advantage that is easy to miss. Cost per outcome is not a new metric that has to be explained to the board. It is cost-income ratio with a narrower denominator. If an agent completes a task for a pound where the fully loaded human cost is thirty, that is not an AI statistic, it is a cost-income argument, expressed in the language the institution is already judged in. The organisations that get funding for this will be the ones that make that translation rather than reporting token counts upwards.
Two related things break at the same time, and both are worth naming.
The first is attribution. A single API key shared across a department produces no chargeback and no way to separate a token spent on somebody's experiment from a token spent on production work that earned something. The Foundation's own practitioner survey names granular AI spend monitoring, tokens, requests and GPU utilisation, as the top tooling request, which is direct evidence that most organisations cannot currently see this.
The second is the recharge model. Technology cost is normally allocated to the business line that consumes it, and that model assumes consumption you can forecast. Agentic consumption is variable, and the person who triggers the spend is frequently not the person who carries it.
The trading book comparison, and what it actually buys you
The argument of this article is that financial services already owns the control framework this problem needs, and is not using it.
Consider what a bank does with any activity that generates variable, real-time exposure. It sets a limit before the activity starts, per desk and per instrument, and the limit is a hard constraint rather than a target. It monitors the position continuously, not at month end. It has an automatic mechanism to stop the activity when the limit is breached, and that mechanism does not depend on someone noticing. It escalates breaches to someone independent of the desk that caused them. And it employs people whose job is watching the limit, structurally separate from the people whose job is using it.
Now consider what most organisations do with an agent. They issue an API key with no cap, to a system whose consumption is a function of its own decisions, monitored through a monthly bill, with no automatic stop, and no separation between the team that builds it and the team that would be responsible for noticing a problem.
A firm that would never let a trader work without a limit will hand an agent an unlimited mandate, because it has classified the spend as technology rather than as exposure. That classification is the mistake. Every one of the five controls above maps cleanly onto agentic AI: a per-workflow token budget, real-time consumption monitoring, an automatic circuit breaker, breach escalation, and independent oversight of the limit. None of it requires new thinking. It requires recognising the shape of the problem.
This is an argument by analogy rather than a claim about regulation, and the distinction matters, which brings us to the point where the article has to be careful.
What regulators are and are not saying
It would be convenient to claim that regulators are demanding cost governance for AI. They are not. Having read across the FCA, the PRA, the Bank of England, ESMA and Parliament, the consistent finding is that AI third-party risk is framed as an operational resilience and concentration question, never as a spend control question. The Bank of England's April 2026 letter responding to the Treasury Committee does not mention cost at all.
What regulators are pressing on is dependency, and that pressure produces the same visibility requirement by a different route.
The Treasury Committee's report on artificial intelligence in financial services, published in January 2026, was unusually direct. It found that "the Financial Conduct Authority, the Bank of England and HM Treasury are not doing enough to manage the risks presented by AI", and that "by taking a wait-and-see approach to AI in financial services, the three authorities are exposing consumers and the financial system to potentially serious harm". On dependency it was blunter still: "UK financial services firms are overly reliant on a small number of US technology firms for AI and cloud services, threatening the sector's operational resilience." It also found that some 75% of UK financial services firms are now using AI, with the heaviest take-up among insurers and international banks operating in the UK.
The machinery to act on this already exists. The UK's Critical Third Parties regime gives the Treasury power to designate providers whose failure could threaten the stability of the financial system, and in July 2026 it made its first four designations: Amazon Web Services EMEA SARL, Google Cloud EMEA Limited, Microsoft Ireland Operations Ltd and Oracle Corporation UK Limited, with regulatory oversight beginning on 13 July 2026. It is worth being precise about what that does and does not mean. Those four are cloud and technology providers, not AI model providers as such, although all four are major AI infrastructure suppliers. The Treasury Committee recommended that major AI providers be designated by the end of 2026. That has not happened yet.
The Bank of England's Financial Stability Report in July 2026 made the concentration point in its own terms: "Where multiple firms rely on the same providers, software components or essential services, a vulnerability at a common supplier could affect several institutions at once." Sarah Breeden's letter to the Committee in April 2026 confirmed that the Bank's AI work includes "concentration risks including from third-party model providers".
The most precise measurement of that dependency comes from ESMA, which surveyed AI adoption across EU securities markets and reported in February 2026 that 62% of respondents rely exclusively on commercial cloud solutions, 41% of them on a single provider. Microsoft was the top provider for almost half of respondents, ahead of OpenAI at 20% and AWS at 8%. Only about 8% of the third-party providers named were domiciled in the EU.
The honest connection between all this and cost governance is not that one requires the other. It is that they demand the same instrumentation. A firm that can answer "what did this workflow cost, and what did it return" can also answer "which provider are we dependent on, for what, and what happens if their price or their availability changes". The visibility is the same visibility. Cost is simply the version of the question that gets funded, because it arrives with a number attached.
What to actually do
Five things follow, none of which require waiting for a standard.
Instrument before you scale. If you cannot attribute token spend to a team, a workflow and a purpose, you cannot manage any of this, and no framework will help. This is the boring prerequisite and it is where most organisations are stuck.
Set the limit as a control, not a target. A per-workflow budget with an automatic stop, enforced in the system rather than in a policy document. If the idea of an agent halting mid-task is unacceptable, that is a useful thing to discover in a design review rather than in an incident.
Measure cost per outcome, and translate it. Report the cost of a completed unit of work against the alternative, in the institution's own financial language, rather than reporting consumption.
Check what your adoption incentives reward. Uber's leaderboard is the cautionary example. Encouraging usage produces usage.
Treat provider concentration as a live risk. Pricing can move in either direction, rate limits can be imposed, and the regulatory direction of travel on critical third parties is clear even though AI model providers have not yet been designated. Knowing what you would do if a provider changed terms mid-year is cheaper than finding out.
The underlying point is simple enough. This is not a new class of problem. It is a familiar class of problem that has been filed in the wrong drawer, and the cost of the misfiling is that the controls a bank already knows how to build are not being built.
References
FinOps as a discipline
FinOps Foundation - What is FinOps? (updated March 2026) https://finops.org/introduction/what-is-finops/
FinOps Foundation - About the FinOps Foundation https://www.finops.org/about/
Vasilio Markanastasakis - The 2026 FinOps Framework, FinOps Foundation (19 March 2026) https://www.finops.org/insights/2026-finops-framework/
FinOps Foundation - FinOps for AI, technology category page https://www.finops.org/framework/technology-categories/ai/
FinOps Foundation - FinOps for AI Overview working group (updated 17 February 2026) https://www.finops.org/wg/finops-for-ai-overview/
FinOps Foundation - State of FinOps 2026 (1,192 respondents, published February 2026) http://data.finops.org/
Tokenomics
Andrew Nhem - FinOps X 2026 Day 1 keynote recap, FinOps Foundation (9 June 2026), source of the J.R. Storment quotations https://www.finops.org/insights/finops-x-2026-day-1-keynote/
The Linux Foundation - Linux Foundation announces the intent to launch the Tokenomics Foundation (3 June 2026) https://www.linuxfoundation.org/press/linux-foundation-announces-the-intent-to-launch-the-tokenomics-foundation-to-establish-open-standards-for-ai-cost-management
What agents cost
Anthropic - How we built our multi-agent research system (13 June 2025), source of the 4x and 15x token figures https://www.anthropic.com/engineering/multi-agent-research-system
Gartner - Gartner Predicts AI Inference Costs Per Agentic Workflow Will Increase More Than Fivefold Through 2028 (17 August 2026) https://www.gartner.com/en/newsroom/press-releases/2026-08-17-gartner-predicts-ai-inference-costs-per-agentic-workflow-will-increase-more-than-fivefold-through-2028
Jim Rowan, Nitin Mittal, Parth Patwari and Ed Burns - The agentic reality check: Preparing for a silicon-based workforce, Deloitte Tech Trends 2026 (10 December 2025) https://www.deloitte.com/us/en/insights/topics/technology-management/tech-trends/2026/agentic-ai-strategy.html
Jake Angelo - Uber's COO on burning through the AI budget, Fortune (26 May 2026) https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-claude-code/
Financial services, dependency and regulation
House of Commons Treasury Committee - Artificial intelligence in financial services, Fifteenth Report of Session 2024-26 (20 January 2026) https://publications.parliament.uk/pa/cm5901/cmselect/cmtreasy/684/report.html
Bank of England - UK financial regulators to begin overseeing Critical Third Parties announced by HMT (10 July 2026) https://www.bankofengland.co.uk/news/2026/july/uk-financial-regulators-to-begin-overseeing-critical-third-parties-announced-by-hmt
FCA - Critical third parties: strengthening UK financial services https://www.fca.org.uk/firms/critical-third-parties-strengthening-uk-financial-services
Sarah Breeden, Bank of England - Response to the Treasury Committee inquiry report on AI in financial services (1 April 2026) https://www.bankofengland.co.uk/-/media/boe/files/letter/2026/response-to-tsc-inquiry-report-on-ai-in-financial-services
Bank of England - Financial Stability Report, July 2026 (7 July 2026) https://www.bankofengland.co.uk/financial-stability-report/2026/july-2026
ESMA - AI adoption and trends in securities markets, TRV Risk Analysis ESMA50-481369926-30599 (20 February 2026) https://www.esma.europa.eu/sites/default/files/2026-02/ESMA50-481369926-30599_TRV_Risk_Analysis_AI_adoption_and_trends_in_securities_markets.pdf
Image
Image by Stevebidmead on Pixabay (Pixabay image ID 814679), used under the Pixabay Content License, which permits free use without attribution. Credit given as a courtesy.

