AI-Assisted Model Validation for Financial Institutions

ALM model validations, with the evidence pass already done.

VetALM puts an AI reader on the whole document set. It extracts every figure with a citation to its exact cell or paragraph, benchmarks the institution's reported interest-rate risk against an independent challenger model, and drafts the workpapers. Your reviewers do what only they can do: exercise judgment and sign off.

  • Every finding cites the exact page, cell, or paragraph
  • Nothing is finalized without a reviewer's sign-off
  • Append-only audit trail, sealed at sign-off
vetalm.com / engagements / 2026 ALM Review
The VetALM findings review screen: a list of detected issues, each with a severity badge, an explanation, the rule that produced it, source citations, and a regulatory mapping.

Findings map to the expectations reviewers and examiners cite

  • SR 11-7
  • OCC 2011-12
  • FDIC FIL-22-2017
  • FFIEC IRR Guidance
  • 2010 Interagency Advisory

Who it's for

Two ways to use the same engine.

An ALM model validation asks whether an institution's interest-rate-risk model is conceptually sound, correctly implemented, and well governed. Whether you're the independent party answering that question or the institution being asked it, the evidence work is identical. It is also the part that eats the calendar.

Primary

Validation & audit firms

You perform independent validations for financial-institution clients. Your constraint isn't demand. It's senior reviewer hours.

  • Issue the document request from a built-in 44-point checklist, and track what arrived, what's still outstanding, and what the institution simply doesn't keep
  • Arrive at the first review meeting with the corpus already read and cited
  • Benchmark reported EVE and NII without building a model per client
  • Export workpapers, an executive summary, and an issues register in your own review flow
How firms use VetALM
Also for

Financial institutions

You own the ALM model, or the model risk function that governs it. A validation or exam is coming, and you'd rather find the gaps yourself.

  • Run the same checks your validator will run, on your own timeline
  • See which of the 44 expected documents you can't currently produce, and decide which of those actually matter
  • Catch stale studies, unsupported overrides, and limit breaches before they're findings
  • Show a documented, evidence-backed self-assessment when the examiner asks
How institutions use VetALM

The platform

Every screen is the working paper.

VetALM isn't a chat window bolted onto a document pile. It is the validation workspace itself: request list, evidence table, findings register, benchmark, and audit trail. AI makes the first pass through every one of them, and hands you the result in the form your workpapers already take.

1 Collect

The document request list is built in.

Forty-four expected documents across seven categories: governance, model documentation, assumptions, results, validation evidence, and source data. Upload into a slot to declare its type, or drop the files in and let the AI classifier read each one and place it. Almost no institution has all forty-four, and a complete set was never the goal. Knowing which ones are missing, and which of those the institution never produced in the first place, is.

  • Coverage is visible at a glance: what's collected, what's still missing
  • The gaps that do matter become findings in their own right
  • PDF, Excel, and Word, with OCR for scans
The Documents tab, showing expected document slots grouped by category with Extracted or Missing status on each.
Nine of forty-four slots filled. A full set is rare, and the useful question is which of the thirty-five outstanding items the validation actually needs.

2 Extract

Every number carries its evidence.

The AI reads betas, decay rates and average lives, EVE and NII shock results, policy limits, governance dates, and back-testing variances out of documents that were never built to be machine-readable. Every extracted field arrives with a confidence score and a locator pointing at the exact spreadsheet cell or document paragraph it came from.

  • Excel and Word are parsed deterministically, so citations resolve to real cells
  • Low-confidence values are routed to a reviewer instead of being assumed
  • Reviewers can correct any value; the correction is recorded, not overwritten
The Extractions tab: a table of extracted fields with values, confidence bars, source evidence locators, and review status.
Ninety-seven fields from ten documents, each traceable back to where it was written.

3 Review

Findings you can argue with.

A hybrid engine runs deterministic rules first, then turns a language model loose on the whole corpus at once. Every finding carries a severity, a plain explanation of why it matters, the citations behind it, and a mapping to the regulatory expectation it speaks to.

  • Approve, edit, or reject each one; rejections require a reason
  • Cross-document contradictions surface that no single-file review would catch
  • The engine refuses to raise a finding it cannot cite
The Findings tab: detected issues with severity badges, explanations, rule identifiers, source citations, and regulatory mappings.
A surge-balance haircut recommended at 25% but configured at 0%, found by reading three documents against each other.

4 Deliver

Workpapers, not a transcript.

When the review is done, the lead signs off and the engagement generates the deliverables in Word, PDF, and Excel: an executive summary, the validation workpaper, and an issues register. The AI drafts the prose; the approved record supplies every number in it.

  • Only reviewer-approved findings reach the deliverables
  • The challenger benchmark and its method note are rendered into the workpaper
  • Sign-off seals the audit trail for the engagement
The Outputs tab, listing a generated issues register, validation workpaper, and executive summary available in Word, PDF, and Excel.
Generated from the approved record. Regenerate, and the audit trail records that too.

Challenger model

The benchmark you'd build by hand, already built.

Supervisory guidance asks a validation to benchmark the model against an independent alternative. In practice that step is often thinned to a peer comparison, because building a second ALM model per engagement isn't realistic. VetALM ships one.

An independent engine, not a reproduction.

VetALM's challenger consumes the account-category rollups that ALM report appendices actually disclose, projects monthly cash flows, and produces economic value of equity and twelve-month net interest income across the full parallel-shock set. It is the platform's own model, deliberately not a re-implementation of the institution's. The AI supplies its inputs by reading the appendix; the model itself is ordinary financial mathematics, and it computes the same answer every time.

  • EVE discounted on the shocked curve; non-maturity deposit coupons follow the shock by their beta, which is what creates deposit franchise value
  • NII on a static balance sheet, with runoff reinvested at market and spreads held constant, so repricing risk is isolated from growth assumptions
  • Variance past tolerance raises a finding; a disagreement on the direction of exposure is escalated regardless of size
  • Runs are append-only and snapshot their own inputs, so the benchmark behind any finding stays reproducible
7Parallel shock scenarios, ±100 to ±400 bps
EVE + NIIBoth metrics benchmarked per scenario
±5 ptsDefault variance tolerance, per engagement
The challenger model workbench: tables comparing challenger EVE and NII against the institution's reported results for each rate shock, with variances flagged and direction disagreements marked as opposed.
Where the challenger and the reported table disagree on sign, the row is marked opposed, and the finding writes itself.

How it works

One auditable pipeline, with a person in the middle of it.

AI does the classifying, the extracting, the reading across documents, and the drafting. Human-in-the-loop isn't a setting on top of that; it is a hard control. The pipeline will not produce a deliverable that a reviewer hasn't approved.

1Ingest
2Classify
3Parse
4Extract
5Detect
6Benchmark
7Human review
8Sign-off & output

Documents are immutable

Every upload is hashed and stored as received. The original is always recoverable, and duplicate uploads are detected by content.

Roles are enforced server-side

Engagement lead, senior reviewer, reviewer, read-only. Who may approve a finding or sign off an engagement is a capability check, not a hidden button.

The trail is append-only

Every action is recorded with an actor, a timestamp, and before/after hashes. Entries are never edited, so tampering is detectable.

The audit trail screen: a table of timestamped actions with the acting entity and before and after content hashes.
The audit trail for a completed engagement, sealed at sign-off.

Detection engine

Twelve rules, then a language model.

Deterministic checks run first: fixed comparison logic, thresholds pulled from the engagement's own extracted policy rather than hard-coded into the product, and a plain explanation attached to every hit. A language model then reads across the full corpus for the contradictions no single rule anticipates, and it is required to cite its evidence or stay quiet.

CheckWhat it catchesTypeDefault severity
Policy limit breachR-LIMIT-BREACH A shock result more adverse than the policy limit for that metric and scenario. DeterministicHigh
Inconsistent governance reportingR-BOARD-INCONSIST A figure reported to ALCO or the board that doesn't match the corresponding model output. HybridHigh
Unsupported overrideR-OVERRIDE-UNSUP A model input set to an override value with no documented rationale behind it. DeterministicHigh
Stale governance reviewR-STALE-REVIEW A policy approval, deposit study, or prior validation past its required refresh cadence. DeterministicHigh
Repeat findingR-REPEAT-FINDING A current condition matching a prior-period finding that is still open or recurring. HybridHigh
Missing assumption supportR-MISSING-SUPPORT An assumption with no source reference or supporting evidence behind it. HybridMedium
Missing back-testingR-MISSING-BACKTEST A metric or period with no back-testing evidence supplied at all. DeterministicMedium
Unexplained back-test varianceR-BACKTEST-VAR A back-test variance beyond the stated tolerance with no explanation recorded. DeterministicMedium
Limit monitoring control gapR-LIMIT-MODULE-OFF Policy limits that are not actually loaded into or monitored by the model. DeterministicMedium
Challenger model varianceR-CHALLENGER-VAR Reported EVE or NII that diverges from the independent benchmark past tolerance, or disagrees with it on direction. DeterministicMedium
Missing expected contentR-MISSING-EXPECTED An expected document or data element absent from the collected corpus. DeterministicMedium
Scenario coverage gapR-SCENARIO-GAP A policy-required shock scenario missing from the shock report. DeterministicLow

Each rule is independently testable and explainable. Where a rule is marked hybrid, the AI supplies the reading and the rule supplies the comparison, so the conclusion stays checkable even though the evidence came out of prose.

Security, privacy & compliance

Built for the people who audit software like this.

Model documentation is confidential, and your clients will ask how it's handled. VetALM holds itself to the model-risk controls it helps you check.

Strict tenant isolation

Every business record carries a tenant, and every query is tenant-scoped in the service layer. Cross-tenant reads aren't a permission. They're absent.

Encrypted in transit and at rest

TLS everywhere, managed keys protecting object storage and the database, private blob access with signed URLs.

Zero data retention AI

VetALM sends document content to a commercial AI provider under a zero-retention configuration. Nothing is kept after the response returns, and nothing is used for training.

Immutable audit trail

Append-only, attributable, time-stamped, and carrying before and after hashes. It is sealed when the engagement is signed off.

Sign-off is a hard control

No conclusion is finalized and no deliverable is generated without an authorized reviewer's approval.

No consumer data, by design

VetALM reviews model documentation, not customer records. Access is scoped to engagement membership, and role capability checks are enforced server-side.

Regulatory references are advisory mappings to assist reviewers. VetALM does not render legal or regulatory opinions.

Results

Less time transcribing. More time deciding.

40–80h 10–25hTarget validation effort per engagement
100%Recall of seeded issues on the golden dataset
0False positives on the clean control
100%Findings traceable to source evidence

Detection is scored against a golden validation dataset before every release. The build fails if recall drops below 90% or the clean control produces more than one false positive, so these numbers cannot silently regress.

Questions

The ones that come up first.

Does VetALM replace the validator?

No, and it's built so it can't. The AI does the evidence pass: reading, extracting, cross-checking, benchmarking, and drafting. Every finding it raises sits in a pending state until a qualified reviewer approves, edits, or rejects it, and no deliverable is generated from unapproved findings. The judgment, the conclusion, and the signature stay with a person.

Do you re-run the institution's ALM model?

No. VetALM reviews the outputs the institution reports. Separately, it runs its own independent challenger model from the account-category rollups disclosed in the report appendix. That is the benchmarking step supervisory guidance asks for, and it is deliberately a different model rather than a reproduction of theirs.

What if a document is a scan, or the format is unusual?

Scanned PDFs route through OCR. Digital PDFs are read natively, including their visual layout. Excel and Word are parsed deterministically so that citations resolve to a real cell or paragraph rather than an approximate guess. Where confidence is low, the value is flagged for a reviewer instead of being asserted.

Can we add our own checks?

Yes. The rule catalog is structured so a new check can be added wherever its trigger is expressible in the rule schema, and thresholds come from each engagement's own extracted policy rather than being fixed in the product.

What happens to our clients' documents?

They're stored encrypted, scoped to a single tenant and engagement, and never used to train a model. The AI provider is configured for zero data retention, so content isn't kept after a response returns. See Security for the full posture.

How do we get started?

A demo on a sample bank dataset takes about thirty minutes, and shows the whole path from document request to signed workpaper. If it looks right, the next step is a pilot on one of your own engagements.

See it run on a real validation.

We'll walk through a complete engagement on a sample bank dataset: the document request, the extracted evidence, the challenger benchmark, and every finding traced back to the cell it came from. Then we'll talk about your corpus.

  • 30-minute guided walkthrough, no slideware
  • Tailored to the documents you actually receive
  • We review reported outputs, so no model re-computation is required

Prefer email? Reach us at sales@vetalm.com.