Diligence · Overview

Put AI investment through the same diligence as every other capital decision

Parallax Intel scores banking AI use cases on common terms — the value peers have evidenced, set against the cost to build, run, and govern — so a bank can weigh unlike use cases the way an investment committee weighs any other claim on capital.

The question facing most banks isn't whether to invest in AI — it's where to start, and where not to. The use cases on the table carry different payoffs and different costs to run safely, and without a consistent basis for weighing them, capital tends to follow the most visible use case rather than the most valuable.

The framework supplies that basis. It asks four questions of every use case and answers each on the same scale:

  1. 01

    Value evidenced — Where has value actually been disclosed in comparable deployments, and how does it compare to the alternatives?

  2. 02

    Regulatory intensity — Which regimes apply to this use case, and how heavily do they bear on it?

  3. 03

    Operating model intensity — What is the ongoing burden to run and govern it, and where in the organization does that burden fall?

  4. 04

    Technical complexity — Which AI components does it rely on, and how difficult are they to build and integrate?

The four-dimensional read

Return

Value evidenced what peers have proven

Cost & risk · operational burden

Regulatory intensity
Operating model intensity
Technical complexity

→ Produces a single read

Fund first Fund with controls Low-stakes pilot Defer
The four scores place a use case on a single funding spectrum — from fund-first to defer.

One question measures return; three measure the cost of realizing it. Read together, they separate the use cases that are well-evidenced and light to run from those whose value is thin — or quietly eroded by the burden of running them.

The standard behind every read

  • Disclosure first. Value, operating, and technology characteristics come from what peers have publicly put on the record — not vendor claims or conjecture.
  • Inference is always labeled. Where disclosure stops, a single labeled inference fills the gap — a technical implication, a regulatory expectation, or a domain-typical practice — tagged by type and never presented as fact.
  • Weighted by what matters. Characteristics that drive more burden count for more, applied consistently across every use case. The factors that move a score are visible; the values behind them are not.
  • Comparative, decision-support. A score places a use case against its peers in the analyzed set — a prioritization benchmark, not an institutional performance measure, and not legal advice.

Coverage

28
Banking
functions
31
Legal
instruments
14 U.S. · 17 Canada, UK & EU
14
Operating-model
categories
164
AI
technologies
Current to July 2026.

Behind the four questions sits a continuously maintained body of analysis: 28 banking functions across the whole bank, 14 operating-model categories, 164 mapped AI technologies, and 31 regulatory instruments across four jurisdictions — the 14 U.S. instruments anchoring the read for U.S. institutions — kept current as the rules move, from the state AI statutes (Colorado's ADMT Act, Texas's TRAIGA, Utah's AIPA) to the 2026 supersession of SR 11-7 by SR 26-2.

The four sections below take each question in turn — what it measures, what feeds it, and what a finished read looks like — with AI in treasury forecasting as the glimpse throughout. The full worked example is available as a personalized brief.

01 Value evidenced

AI value is easy to claim and hard to evidence

Every vendor deck and internal champion has reason to round up: pilots get reported as wins, and “AI-powered” gets attached to outcomes it didn't drive. The anchor that holds under scrutiny is the public record — what a peer bank, or the vendor it named, has stated openly, with a number where one exists. And silence is not absence: a benefit no one has disclosed is recorded as missing evidence, never as proof the value isn't real.

The read has two parts:

  • Breadth — whether benefit is disclosed across the kinds of value a bank actually weighs: nine benefit dimensions grouped into four categories — commercial impact, operational efficiency, model and process quality, and risk and compliance.
  • Depth — the single strongest result on the record. Breadth can flatter a use case with many small, well-reported claims; depth surfaces the one most consequential outcome a peer has actually disclosed.

What counts is weighted by how well it's disclosed: a realized benefit over a planned one, a quantified figure over a directional phrase, a named primary source over a passing mention, a bank-wide outcome over a single anecdote.

Evidence breadth · six sample use cases

Scroll horizontally to compare all nine dimensions →

Commercial impact Operational eff. Model & process qual. Risk & compliance
Financial Customer Adoption Labor Speed Quality Depth Risk Compliance
Treasury Forecasting
AML
Credit Underwriting
Fraud Detection
Algorithmic Trading
Pricing Optimization
Significant Moderate Limited None disclosed
Cell shading is strength of evidence disclosed — not value generated — across the six sample use cases from the benefit-claims register.

The glimpse: Treasury Forecasting

The most broadly evidenced of the six sample use cases, scoring 85 against a register of 111 disclosed benefit claims across the cohort. Its evidence concentrates in labor saved and forecasts that reach further, across four major U.S. deployments — Bank of America's CashPro, JPMorgan's Cash Flow Intelligence, PNC's PINACLE, and Citi's Cashforce / TIS. The 85 says its value is the best documented — not that it is the most valuable bet a given bank could make.

Value evidenced · Treasury Forecasting

85 / 100
Value Evidenced Score

Most broadly evidenced of the six sample use cases.

Higher means more comprehensively evidenced across peers — not more valuable in absolute terms.


Strongest disclosed result

~250,000 client hours saved in 2025

across 3,000+ corporate clients · Bank of America CashPro

A well-evidenced use case can still be costly to run safely. The next three sections read those costs — beginning with regulatory intensity.

02 Regulatory intensity

What drives regulatory burden isn't AI sophistication — it's what the use case does

The instinct is to count the laws on the books, or to assume the most advanced AI draws the most scrutiny. Both mislead. Regulatory weight is set by the specific characteristics of each use case: what the workflow does with its outputs, which obligations that workflow actually triggers, and how the exposure enforces. The sharpest single line is whether the system reaches decisions about people — their credit, their eligibility, their treatment — but it is one line among several.

For each use case, the framework builds a regulatory surface by asking three things of all 31 instruments — nine regulatory domains, four jurisdictions, anchored on the 14 U.S. instruments:

  • Does it apply? Applicability follows what the workflow does — not where the bank operates, and not the mere presence of AI.
  • How directly does it bear? Each applicable instrument is rated High, Medium, Low, or Not relevant by how central its obligations are to the use case's typical workflow.
  • How hard would it bite? Enforcement character — supervisory or statutory, penalty type and severity — and jurisdictional reach, including whether obligations extend across borders.

Relevance is scored separately from severity. Keeping the two apart stops a heavy-penalty law a use case barely touches from masquerading as a central obligation — the most common way regulatory heat maps mislead.

U.S. regulatory relevance · Treasury Forecasting

Regime familyInstrument(s)Relevance
Third-party & outsourcing risk
Interagency TPRM guidance · primary driver
High
Model risk
SR 26-2
Medium
Tech governance
FFIEC IT Examination Handbook · OCC Heightened Standards
Low
Financial crime & payment systems
BSA / AML
Not relevant
AI-specific state laws
Colorado ADMT Act (SB 26-189) · Texas TRAIGA · Utah AIPA
Not relevant
Consumer protection & fair lending
FCRA/Reg V · ECOA/Reg B · UDAAP
Not relevant
Data privacy & security
GLBA · CCPA/CPRA (+ CPPA ADMT) · Illinois BIPA
Not relevant
Relevance to the typical workflow, rated separately from severity and ordered most-relevant first; burgundy marks the single High row. Coverage spans 31 instruments across four jurisdictions; the 14 U.S. instruments are shown. Decision-support — counsel required for definitive applicability and compliance determinations.

The glimpse: Treasury Forecasting

The lightest regulatory profile of the six sample use cases, at 49 — and lighter in kind, not only in degree. Because it produces analysis rather than decisions about people, consumer-protection, fair-lending, and the state AI statutes read as Not relevant. What remains concentrates in three supervisory regimes — third-party risk, model risk (SR 26-2), and tech governance — and runs through examination, not statute.

Regulatory intensity · Treasury Forecasting

49 / 100
Regulatory Intensity Score

The lightest profile of the six sample use cases.

Higher means heavier regulatory intensity on the typical workflow — not a measure of a bank's compliance, and not legal advice.


Character of exposure

Supervisory examination, not statutory penalty

Consumer-protection and fair-lending law does not engage

Knowing which rules bite is one cost; running and governing the system that satisfies them is another — and it is permanent.

03 Operating model intensity

The build is a one-time cost; running and governing an AI system is permanent — and it's where realized value quietly leaks away

This is the question most business cases under-price: what does it cost to run and govern this system, day after day — and where in the bank does that cost land? The build is the cheap part. Running is continuous and cross-functional — drawing steadily on model risk, data operations, and governance — and it is where evidenced value erodes when the intensity was never priced in. That burden is unevenly distributed, often shaped by what regulation and the technology demand rather than freely chosen; knowing where it concentrates is what lets a bank budget for it rather than discover it after the system is live.

The framework characterizes how a system is actually built, run, and governed across 14 operating-model categories in four layers:

  • Algorithmic design — model architecture, and the learning paradigm.
  • Data and feature lifecycle — feature engineering, training-data scope, and retraining cadence.
  • Operational behavior — inference mode, human-in-the-loop controls, exception handling, and how outputs feed downstream decisions.
  • Governance and trust — explainability, fairness controls, model validation, data-quality controls, and data privacy.
Operating-model intensity · six sample use cases

Scroll horizontally to compare all six sample use cases →

Algorithmic Trading AML Credit Underwriting Fraud Detection Pricing Optimization Treasury Forecasting
Algorithmic design
Model architecture
Paradigm type
Data & feature lifecycle
Feature engineering
Training data
Model refresh
Operational behavior
Inference mode
Human-in-the-loop
Exception handling
Output integration
Governance & trust
Explainability
Fairness & discrimination
Model validation
Data controls
Data privacy
High Moderate Limited Low / N/A
Intensity of the operating-model choice for each category, across the six sample use cases from the Functional Operating Model table.

Banks rarely publish their full operating model — so this is where the labeled-inference method does its heaviest work. Each category is read from the public record first; where the record stops, one labeled inference completes it, so a reader can see exactly where disclosure ends and inference begins.

Reading the operating model · JPM Cash Flow Intelligence

CategoryCharacterization · basisDetermined by
Human-in-the-loop
Pre-decision review

Liquidity decisions remain with human teams

Disclosed
Output integration
Informational / analytical

Dashboards and forecasts, not automated execution

Disclosed
Model validation
Independent formal validation

SR 26-2 expects it for a material model

Regulatory expectation
Retraining
Calendar-scheduled

Enterprise treasury forecasting retrains on cycle

Domain-typical
Data controls
Automated validation

Automated reconciliation implies validation at ingestion

Technical implication
Burgundy marks what the bank has disclosed; the muted tags mark where a single labeled inference fills what the record leaves open. Five categories shown — chosen characterizations only, not option sets or scores.

The glimpse: Treasury Forecasting

The lightest operating profile of the six, at 56 — but “light overall” hides where the weight sits. Most of its run-and-govern choices sit at the cohort floor; the intensity concentrates in one category — data sourcing, the multi-source integration across client ERP systems, external bank accounts, and internal systems that the forecasts depend on — with model validation and data privacy shaped by regulation rather than choice. Even a light use case carries a specific, locatable burden, and the value of the read is knowing where it sits before the system is live.

Operating model intensity · Treasury Forecasting

56 / 100
Operating Model Intensity Score

The lightest profile of the six sample use cases.

Higher means a heavier ongoing burden to run and govern — not a measure of a bank's operating maturity.


Where the weight sits

Multi-source data integration

Light to run in most categories; the burden concentrates in data sourcing

Running the system is one cost; the difficulty of building it in the first place is another. The final section reads technical complexity.

04 Technical complexity

Build difficulty isn't one thing — frontier AI is often cheaper to build than conventional models wired into hard systems

The temptation is to read “AI complexity” as a single dial: the more advanced the model, the harder the build. It doesn't hold. For a great deal of banking AI the model is the cheapest part — the real cost sits in the data pipeline that feeds it, or in the systems of record it has to integrate with. The framework decomposes build difficulty into its sources, reading three components:

  • AI sophistication — which technologies the system uses, mapped against a taxonomy of 164 AI technologies in three layers: core AI, applied AI, and lifecycle and infrastructure.
  • Data complexity — how many and which data modalities it works across, from structured tables and time series through graph relationships, text, and beyond.
  • Integration depth — how many systems the model must connect to, and how deep those connections run.

The read records not only what a system uses but what it notably does not — the absence of large language models, graph networks, or unstructured data is part of the profile, and recording it stops a conventional system from being scored as if it were frontier.

AI components · Treasury Forecasting

AI components

Disclosed

Ensemble methods

Time-series forecasting

Scenario simulation

AutoML

Drift monitoring

Inferred

Neural networks (domain-typical)

Anomaly detection (technical implication)

Notably absent

Large language models

Generative AI

Graph neural networks

Federated learning

Computer vision

NLP

Data modalities

Disclosed

Structured / tabular

Time-series / temporal

Absent

Graph / network

Text

Image

Video

Voice

Code

Geospatial

Integration

Disclosed

Multi-bank ingestion

ERP (SAP)

TMS

Components from public disclosure; inference labeled by type; absences recorded. Chosen components shown — not the per-component complexity weights or the score.

The glimpse: Treasury Forecasting

The lightest build of the six, at 40 — and the decomposition shows why. Its disclosed AI is conventional and mature — ensemble methods, time-series forecasting, scenario simulation — on a narrow, two-modality data layer. What complexity exists lives almost entirely in the integration: multi-bank ingestion, ERP, and treasury management systems — few system classes, each a system of record that tolerates no error. The weight sits in the plumbing, not the model — the same picture the operating read showed.

Technical complexity · Treasury Forecasting

40 / 100
Tech Complexity Score

The lightest build of the six sample use cases.

Higher means greater build and integration difficulty — not a measure of a bank's engineering capability.


Where the cost sits

Integration, not the model

Conventional AI on a narrow data layer

Diligence · Conclusion

Where AI capital goes first

That completes the read: value, set against the three costs of realizing it. Placed on common terms, the well-evidenced, light-to-run cases earn capital first — and the ones whose value is outweighed by the cost of building, running, and governing them can wait, on evidence rather than instinct. Producing that placement is what the framework is for.

Treasury Forecasting · the read on common terms

85
Value
evidenced
49
Regulatory
intensity
56
Operating model
intensity
40
Technical
complexity

→ Placed on the funding spectrum

Fund first Fund with controls Low-stakes pilot Defer
High evidenced value against the lightest cost profile of the six — a fund-first placement.
Personalized brief

Read the full worked example

The complete Treasury Forecasting brief carries what this page only glimpses: the full four-dimension read, every characterization with its evidence or labeled inference, and the peer-by-peer detail across Bank of America, JPMorgan, PNC, and Citi. Issued as a personalized, named-recipient document.

Deciding where to start with AI — or where to take it next?

A structured, evidence-based read may be useful to how those decisions get made.

Start a conversation