The Two Ledgers

The Only AI Cost That Gets Cheaper

Token sticker shock is real. It is also the only part of a bank's AI cost base riding a deflation curve. The ledger below is unmetered, priced in labor, and compounding.

Goldman Sachs's chief information officer spent part of this spring warning that CFOs should brace for "a token sticker shock," with bills arriving that nobody modeled. He is already late to some inboxes. One bank told Evident Insights that its token costs had grown 250% since the start of the year. Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027, because costs escalate, business value stays unclear, or risk controls fall short. Read that list again: only one of the three reasons is about the meter.

The interesting question is not whether the bill is real. It is why the shock is coming from the one input in the entire program that is getting cheaper.

The unit economics run in one direction. Stanford's 2025 AI Index found that querying a model at the level of GPT-3.5 cost about $20 per million tokens in November 2022 and about seven cents by October 2024, a fall of more than 280-fold in under two years. Goldman Sachs Research expects per-token inference costs to keep falling 60% to 70% a year at the hardware level. Few other lines in a bank's cost base behave like this.

Total spend runs in the other direction. Menlo Ventures puts enterprise spending on generative AI at $37 billion in 2025, roughly triple the prior year, and states the paradox plainly: net spend keeps rising even as the cost of inference falls, because usage grows by orders of magnitude. Agentic workflows consume tokens the way chat never did, and frontier reasoning models cost more per token while burning more tokens per task. Cheaper units, multiplied by exploding volume and richer models, still produce a bigger bill.

Little of this is mismanagement. It is what a deflating input looks like when demand is elastic. The shock is not evidence that tokens are expensive. It is evidence that tokens are the only cost anyone is watching.

Banks are learning to watch them well. A discipline is forming around the meter: token budgets, routing traffic to cheaper models, caching, usage caps, and dedicated inference optimization teams. It is the cloud cost playbook applied to AI, and the practitioners who manage cloud spend now describe AI as the fastest-growing line they oversee.

Two things about the meter deserve more attention than they get. First, metered is not the same as measured where decisions are made. Tokens are billed at the vendor account and the platform, not at the use case. A shared assistant serves a dozen functions on one invoice, and agents call other agents. Attributing consumption to a specific use case takes tagging and chargeback design that few banks have built. The cost is metered at the invoice and unallocated at the use case.

Second, the meter is the half banks can see at all. Evident's analysis of the 50 largest banks finds a fast-growing minority now publishing a headline figure for AI's benefit, mostly forward-looking targets rather than realized returns. Cost figures appear at neither level.

Underneath the meter sits a second ledger, and it arrives without an invoice. It is the independent validation a model summons before and after it goes live, the people who oversee its outputs, the documentation that exists because an examiner may ask for it, the conformity work when a system touches a regulated decision, the vendor risk lifecycle behind every model provider, and the process redesign that makes any of it pay off. Evident adds the quietest entry: until workflows actually change, a bank pays for the old process and for the AI layered on top of it at the same time.

Two things price this ledger: labor and regulatory volume. Neither is deflating. U.S. federal agencies introduced 59 AI-related regulations in 2024, more than double the year before.

How large is the governance side? The only such figure I can assemble entirely from primary sources is European. The European Commission's own impact-assessment line items, modeled in 2021, sum to roughly €104,000 to €108,000 a year of governance run-rate for a provider operating a single high-risk system: about €29,000 in compliance labor, about €71,000 in quality-management oversight, plus periodic conformity fees. Spread a quality system across ten models and the per-system figure falls toward €37,000 to €44,000. Those are modeled numbers, for providers, in one jurisdiction. And that is precisely the point: there is no comparable U.S. figure to cite, because banks rarely disclose what a model costs to validate and oversee under supervisory model risk expectations. The governance ledger is not just unmetered. It is undisclosed.

Rules of thumb gesture at the ratio. BCG's guideline allocates 10% of an AI transformation's effort to algorithms, 20% to technology and data, and 70% to people and process. McKinsey suggests planning three dollars of change management for every dollar of model development. These are rules, not measurements. They point the same way.

Here is the asymmetry worth writing on a whiteboard. Among the recurring cost components of a bank's AI program, the token is the only one whose unit price sits on a structural deflation curve. Every governance component, from validation and oversight to documentation and conformity, is priced in labor and indexed to regulatory volume, and both are rising.

The corollary explains a great deal of frustration. Nearly every optimization lever a bank owns works on the deflating half. Routing, caching, compression, and caps all squeeze the cost that was falling anyway. Almost none of them touches the half that compounds.

It gets worse, because the levers interact. Several of the standard ways to cut the token bill grow the governance bill.

Fine-tuning is the cleanest example. A bank that adapts a third-party model on its own data to cut inference costs can change what it answers for: modify a model deeply enough and the bank starts to be treated less like a buyer and more like a builder, and a heavier apparatus attends the system from then on. An engineering decision made to save variable cost books a step-change in fixed cost.

Model routing is a second. Diversifying across vendors to chase price summons the full third-party risk lifecycle for each new relationship: due diligence, contracting, ongoing monitoring, and exit planning. The routing table gets cheaper per token and more expensive per vendor.

And every additional model, added for any reason, enters the inventory and adds a permanent increment to the run-rate. Governance expansion is cumulative rather than linear. Costs arrive with each model and rarely leave with one.

Then there is the multiplier no engineering choice controls. One regulatory change can touch every affected system in a portfolio at once. Colorado's revised artificial intelligence statute removes the exemption for financial institutions effective January 2027; on that date, a single legislative act reprices governance across a bank's entire AI book simultaneously. Tokens do not do that.

The two ledgers also scale in opposite directions, and that lands hardest on regional banks. Governance cost per model falls as a portfolio grows: frameworks amortize, and a quality system, a validation function, and an inventory process spread across every model that follows the first. Token cost per model does not fall with portfolio size in the same way; the meter charges the tenth model the same rate as the first.

Now consider a typical mid-sized bank. Supervisory expectations scale with size, but they scale more slowly than portfolios shrink: a mid-sized bank fields a recognizable version of the same governance shape across a fraction of the use cases and a thinner second line. The regional bank sits at the expensive end of the governance curve while paying the same token prices as everyone else. For that institution, the question is not which model is cheapest to run. It is which use cases carry enough evidenced value to justify the burden that arrives with them, and which do not earn their place at all.

There is a discipline for the metered half now. FinOps for AI will tell a bank what a use case costs to run, or at least what its platforms cost to run. There is not yet a comparable discipline for the other ledger, one that prices what a use case costs to be allowed to run: the validation, the oversight, the documentation, the vendor lifecycle, and the process change, jurisdiction by jurisdiction, before approval rather than after. That gap is where AI programs actually get surprised.

The token bill will keep shocking people, and unit prices will keep falling while it does. Both ledgers belong in the same file before a use case is approved: the one that deflates and the one that compounds. The cheapest AI decision is still the one a bank prices completely, on both ledgers, before saying yes.