No bank would book a loan on the borrower's own slideshow. The credit file needs audited financials, collateral, and a repayment history before a dollar moves. AI purchases, still new enough that no comparable standard has settled into place, often clear committee on lighter evidence: a vendor's case study, an internal champion's conviction, a business case assembled before anyone knew which questions to ask. Gartner finds that 72% of CIOs report breaking even or losing money on AI, and that cost calculations miss by 500% to 1,000% when organizations don't yet know how generative AI costs scale. I've spent the past year reading the public record of AI at major US banks: the results disclosed, the rules that apply, the costs to build and run. The rigor banks bring to every other large purchase simply hasn't had time to reach this one.
The first page of that file is the benefit, and the evidence exists. Take one unglamorous product: AI cash-flow forecasting for corporate clients, run by four major US banks, all of which have published results. Bank of America reports its clients saved more than 250,000 hours in 2025. Realized, quantified, disclosed by the bank itself, bank-wide: about as good as AI evidence gets. Set it beside a software vendor's claim that forecast accuracy jumps from 50% to 80% within months. Same arithmetic, different witness: the seller wrote the number. A vendor's accuracy claim is an audition; a bank's disclosed result is a track record.
Even good numbers reward a close reading of their footnotes. A JPMorgan client reports up to 90% effort reduction, and the fine print scopes that figure to resolving inquiries. One task, not the whole treasury job. "Ninety percent of what" is the first question underwriting asks. User counts invite the same care: 3,000 companies here, "tens of thousands of users" there measure adoption, not that the thing was worth building. Five questions grade any AI-benefit claim: did it happen, who says so, how specific, at what scale, and in what currency?
The second page is the cost, and it is the one the market rarely quotes. Vendors price the build. The run mostly goes unpriced: the model validation, retraining, monitoring, and compliance reviews that continue as long as the system does. The run stays invisible because it is rarely a new hire. It is a tax on existing jobs, in slices of legal, privacy, risk, data, audit, and operations, and a tax spread across existing jobs lands on no single budget line, which is precisely why it is so easy to leave out of the business case.
Nor is the burden fixed; it follows the promise. The forecasting product is cheap to run because of four choices: it claims a forecast rather than a fact, informs rather than executes, runs overnight rather than live, and deals with corporate cash rather than people. If forecasts start moving money, running live, or facing consumers, the bill reprices at once. Those dials rarely appear in a pitch, and the ones that stay hidden are the ones that set the real price.
The third page is the rulebook. The temptation is to wait for AI-specific rules to arrive; the reality is that there is no AI exemption from banking law. Regulation follows the decision, not the technology. AI inherits the rules that already govern whatever it touches, then concentrates them along what the decision can harm: people, markets, or the bank itself. A credit decision answers to nearly every instrument on the shelf; this forecasting tool, deciding nothing about any individual, escapes most of them. An algorithmic trading model touches no individual either, yet carries one of banking's heaviest rulebooks, because its mistakes hit the P&L instantly. This spring's SR 26-2 replaced fifteen years of model-risk guidance and pointedly left generative and agentic AI outside its scope. Fewer prescriptions, not fewer expectations: the guidance is non-binding, yet supervisors can still act on anything unsafe or unsound, so demonstrating sound operation rests on each bank's judgment.
The fourth page is the technology, where the reflex worth resisting is treating sophistication as value. The market rewards impressive-sounding AI; a bank pays for it either way. The best-documented product in this story runs on machinery whose core methods predate the smartphone; the breakthrough was newly reachable client data, not new math. In most bank AI the model works in weeks and the wiring into decades-old systems takes quarters. The clever part is often ordinary, the plumbing is the expense, and the most expensive phrase in any pitch is "state of the art."
The four pages don't often line up: simple builds carry heavy rulebooks, the heaviest rulebooks attach where no individual is touched, and the best proof sits on the lightest burden. Benefits, regulations, complexity, and running cost are different axes, and assuming they travel together is how capital gets misallocated. Across the functions I studied, the strongest documented value case was often the least glamorous product. It isn't a law; it's a reason to look before assuming.
None of this argues for timidity: some of the most valuable territory in banking AI is also the heaviest, and the point is to see the weight before signing, not to avoid it. "Not evidenced in the public record" is itself a finding: proceeding is conviction, sometimes exactly right, and conviction is worth sizing as conviction rather than booking as proof.
The natural owner of this discipline is the one the bank already trusts with commitments of this size: the CFO, the CRO, and the rest of the executive committee, the place where every other large decision already survives an evidence standard. Four questions come before any AI dollar moves. What has actually been proven, and by whom? What rulebook does this decision answer to? What does it really take to build? And what does it cost to run, forever? A bank's cheapest AI decision is the one it makes before buying anything. Banks bring this rigor to everything else they underwrite; the opportunity now is to extend it to their AI.