Put AI investment through the same diligence as every other capital decision
Parallax Intel scores banking AI use cases on common terms — the value peers have evidenced, set against the cost to build, run, and govern — so a bank can weigh unlike use cases the way an investment committee weighs any other claim on capital.
The question facing most banks isn't whether to invest in AI — it's where to start, and where not to. The use cases on the table carry different payoffs and different costs to run safely, and without a consistent basis for weighing them, capital tends to follow the most visible use case rather than the most valuable.
The framework supplies that basis. It asks four questions of every use case and answers each on the same scale:
-
01
Value evidenced — Where has value actually been disclosed in comparable deployments, and how does it compare to the alternatives?
-
02
Regulatory intensity — Which regimes apply to this use case, and how heavily do they bear on it?
-
03
Operating model intensity — What is the ongoing burden to run and govern it, and where in the organization does that burden fall?
-
04
Technical complexity — Which AI components does it rely on, and how difficult are they to build and integrate?
The four-dimensional read
Return
Cost & risk · operational burden
→ Produces a single read
One question measures return; three measure the cost of realizing it. Read together, they separate the use cases that are well-evidenced and light to run from those whose value is thin — or quietly eroded by the burden of running them.
The standard behind every read
- Disclosure first. Value, operating, and technology characteristics come from what peers have publicly put on the record — not vendor claims or conjecture.
- Inference is always labeled. Where disclosure stops, a single labeled inference fills the gap — a technical implication, a regulatory expectation, or a domain-typical practice — tagged by type and never presented as fact.
- Weighted by what matters. Characteristics that drive more burden count for more, applied consistently across every use case. The factors that move a score are visible; the values behind them are not.
- Comparative, decision-support. A score places a use case against its peers in the analyzed set — a prioritization benchmark, not an institutional performance measure, and not legal advice.
Coverage
functions
instruments
categories
technologies
Behind the four questions sits a continuously maintained body of analysis: 28 banking functions across the whole bank, 14 operating-model categories, 164 mapped AI technologies, and 31 regulatory instruments across four jurisdictions — the 14 U.S. instruments anchoring the read for U.S. institutions — kept current as the rules move, from the state AI statutes (Colorado's ADMT Act, Texas's TRAIGA, Utah's AIPA) to the 2026 supersession of SR 11-7 by SR 26-2.
The four sections below take each question in turn — what it measures, what feeds it, and what a finished read looks like — with AI in treasury forecasting as the glimpse throughout. The full worked example is available as a personalized brief.
AI value is easy to claim and hard to evidence
Every vendor deck and internal champion has reason to round up: pilots get reported as wins, and “AI-powered” gets attached to outcomes it didn't drive. The anchor that holds under scrutiny is the public record — what a peer bank, or the vendor it named, has stated openly, with a number where one exists. And silence is not absence: a benefit no one has disclosed is recorded as missing evidence, never as proof the value isn't real.
The read has two parts:
- Breadth — whether benefit is disclosed across the kinds of value a bank actually weighs: nine benefit dimensions grouped into four categories — commercial impact, operational efficiency, model and process quality, and risk and compliance.
- Depth — the single strongest result on the record. Breadth can flatter a use case with many small, well-reported claims; depth surfaces the one most consequential outcome a peer has actually disclosed.
What counts is weighted by how well it's disclosed: a realized benefit over a planned one, a quantified figure over a directional phrase, a named primary source over a passing mention, a bank-wide outcome over a single anecdote.
Scroll horizontally to compare all nine dimensions →
| Commercial impact | Operational eff. | Model & process qual. | Risk & compliance | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Financial | Customer | Adoption | Labor | Speed | Quality | Depth | Risk | Compliance | |
| Treasury Forecasting | — | — | — | ||||||
| AML | — | — | — | — | |||||
| Credit Underwriting | — | — | |||||||
| Fraud Detection | — | — | — | — | — | ||||
| Algorithmic Trading | — | — | — | — | — | — | |||
| Pricing Optimization | — | — | — | — | — | — | — | ||
The glimpse: Treasury Forecasting
The most broadly evidenced of the six sample use cases, scoring 85 against a register of 111 disclosed benefit claims across the cohort. Its evidence concentrates in labor saved and forecasts that reach further, across four major U.S. deployments — Bank of America's CashPro, JPMorgan's Cash Flow Intelligence, PNC's PINACLE, and Citi's Cashforce / TIS. The 85 says its value is the best documented — not that it is the most valuable bet a given bank could make.
Value evidenced · Treasury Forecasting
Most broadly evidenced of the six sample use cases.
Higher means more comprehensively evidenced across peers — not more valuable in absolute terms.
Strongest disclosed result
~250,000 client hours saved in 2025
A well-evidenced use case can still be costly to run safely. The next three sections read those costs — beginning with regulatory intensity.
What drives regulatory burden isn't AI sophistication — it's what the use case does
The instinct is to count the laws on the books, or to assume the most advanced AI draws the most scrutiny. Both mislead. Regulatory weight is set by the specific characteristics of each use case: what the workflow does with its outputs, which obligations that workflow actually triggers, and how the exposure enforces. The sharpest single line is whether the system reaches decisions about people — their credit, their eligibility, their treatment — but it is one line among several.
For each use case, the framework builds a regulatory surface by asking three things of all 31 instruments — nine regulatory domains, four jurisdictions, anchored on the 14 U.S. instruments:
- Does it apply? Applicability follows what the workflow does — not where the bank operates, and not the mere presence of AI.
- How directly does it bear? Each applicable instrument is rated High, Medium, Low, or Not relevant by how central its obligations are to the use case's typical workflow.
- How hard would it bite? Enforcement character — supervisory or statutory, penalty type and severity — and jurisdictional reach, including whether obligations extend across borders.
Relevance is scored separately from severity. Keeping the two apart stops a heavy-penalty law a use case barely touches from masquerading as a central obligation — the most common way regulatory heat maps mislead.
U.S. regulatory relevance · Treasury Forecasting
The glimpse: Treasury Forecasting
The lightest regulatory profile of the six sample use cases, at 49 — and lighter in kind, not only in degree. Because it produces analysis rather than decisions about people, consumer-protection, fair-lending, and the state AI statutes read as Not relevant. What remains concentrates in three supervisory regimes — third-party risk, model risk (SR 26-2), and tech governance — and runs through examination, not statute.
Regulatory intensity · Treasury Forecasting
The lightest profile of the six sample use cases.
Higher means heavier regulatory intensity on the typical workflow — not a measure of a bank's compliance, and not legal advice.
Character of exposure
Supervisory examination, not statutory penalty
Knowing which rules bite is one cost; running and governing the system that satisfies them is another — and it is permanent.
The build is a one-time cost; running and governing an AI system is permanent — and it's where realized value quietly leaks away
This is the question most business cases under-price: what does it cost to run and govern this system, day after day — and where in the bank does that cost land? The build is the cheap part. Running is continuous and cross-functional — drawing steadily on model risk, data operations, and governance — and it is where evidenced value erodes when the intensity was never priced in. That burden is unevenly distributed, often shaped by what regulation and the technology demand rather than freely chosen; knowing where it concentrates is what lets a bank budget for it rather than discover it after the system is live.
The framework characterizes how a system is actually built, run, and governed across 14 operating-model categories in four layers:
- Algorithmic design — model architecture, and the learning paradigm.
- Data and feature lifecycle — feature engineering, training-data scope, and retraining cadence.
- Operational behavior — inference mode, human-in-the-loop controls, exception handling, and how outputs feed downstream decisions.
- Governance and trust — explainability, fairness controls, model validation, data-quality controls, and data privacy.
Scroll horizontally to compare all six sample use cases →
| Algorithmic Trading | AML | Credit Underwriting | Fraud Detection | Pricing Optimization | Treasury Forecasting | |
|---|---|---|---|---|---|---|
| Algorithmic design | ||||||
| Model architecture | ||||||
| Paradigm type | ||||||
| Data & feature lifecycle | ||||||
| Feature engineering | ||||||
| Training data | ||||||
| Model refresh | ||||||
| Operational behavior | ||||||
| Inference mode | — | |||||
| Human-in-the-loop | ||||||
| Exception handling | ||||||
| Output integration | — | |||||
| Governance & trust | ||||||
| Explainability | ||||||
| Fairness & discrimination | — | — | ||||
| Model validation | ||||||
| Data controls | ||||||
| Data privacy | — | |||||
Banks rarely publish their full operating model — so this is where the labeled-inference method does its heaviest work. Each category is read from the public record first; where the record stops, one labeled inference completes it, so a reader can see exactly where disclosure ends and inference begins.
Reading the operating model · JPM Cash Flow Intelligence
Liquidity decisions remain with human teams
Dashboards and forecasts, not automated execution
SR 26-2 expects it for a material model
Enterprise treasury forecasting retrains on cycle
Automated reconciliation implies validation at ingestion
The glimpse: Treasury Forecasting
The lightest operating profile of the six, at 56 — but “light overall” hides where the weight sits. Most of its run-and-govern choices sit at the cohort floor; the intensity concentrates in one category — data sourcing, the multi-source integration across client ERP systems, external bank accounts, and internal systems that the forecasts depend on — with model validation and data privacy shaped by regulation rather than choice. Even a light use case carries a specific, locatable burden, and the value of the read is knowing where it sits before the system is live.
Operating model intensity · Treasury Forecasting
The lightest profile of the six sample use cases.
Higher means a heavier ongoing burden to run and govern — not a measure of a bank's operating maturity.
Where the weight sits
Multi-source data integration
Running the system is one cost; the difficulty of building it in the first place is another. The final section reads technical complexity.
Build difficulty isn't one thing — frontier AI is often cheaper to build than conventional models wired into hard systems
The temptation is to read “AI complexity” as a single dial: the more advanced the model, the harder the build. It doesn't hold. For a great deal of banking AI the model is the cheapest part — the real cost sits in the data pipeline that feeds it, or in the systems of record it has to integrate with. The framework decomposes build difficulty into its sources, reading three components:
- AI sophistication — which technologies the system uses, mapped against a taxonomy of 164 AI technologies in three layers: core AI, applied AI, and lifecycle and infrastructure.
- Data complexity — how many and which data modalities it works across, from structured tables and time series through graph relationships, text, and beyond.
- Integration depth — how many systems the model must connect to, and how deep those connections run.
The read records not only what a system uses but what it notably does not — the absence of large language models, graph networks, or unstructured data is part of the profile, and recording it stops a conventional system from being scored as if it were frontier.
AI components · Treasury Forecasting
AI components
Disclosed
Ensemble methods
Time-series forecasting
Scenario simulation
AutoML
Drift monitoring
Inferred
Neural networks (domain-typical)
Anomaly detection (technical implication)
Notably absent
Large language models
Generative AI
Graph neural networks
Federated learning
Computer vision
NLP
Data modalities
Disclosed
Structured / tabular
Time-series / temporal
Absent
Graph / network
Text
Image
Video
Voice
Code
Geospatial
Integration
Disclosed
Multi-bank ingestion
ERP (SAP)
TMS
The glimpse: Treasury Forecasting
The lightest build of the six, at 40 — and the decomposition shows why. Its disclosed AI is conventional and mature — ensemble methods, time-series forecasting, scenario simulation — on a narrow, two-modality data layer. What complexity exists lives almost entirely in the integration: multi-bank ingestion, ERP, and treasury management systems — few system classes, each a system of record that tolerates no error. The weight sits in the plumbing, not the model — the same picture the operating read showed.
Technical complexity · Treasury Forecasting
The lightest build of the six sample use cases.
Higher means greater build and integration difficulty — not a measure of a bank's engineering capability.
Where the cost sits
Integration, not the model
Where AI capital goes first
That completes the read: value, set against the three costs of realizing it. Placed on common terms, the well-evidenced, light-to-run cases earn capital first — and the ones whose value is outweighed by the cost of building, running, and governing them can wait, on evidence rather than instinct. Producing that placement is what the framework is for.
Treasury Forecasting · the read on common terms
evidenced
intensity
intensity
complexity
→ Placed on the funding spectrum
Read the full worked example
The complete Treasury Forecasting brief carries what this page only glimpses: the full four-dimension read, every characterization with its evidence or labeled inference, and the peer-by-peer detail across Bank of America, JPMorgan, PNC, and Citi. Issued as a personalized, named-recipient document.
Deciding where to start with AI — or where to take it next?
A structured, evidence-based read may be useful to how those decisions get made.