Automating 70% of document review with an audit-grade RAG platform
How a regional lender automated the majority of its credit-file document review with a retrieval system built for citation, auditability and on-premises deployment.
Client: Regional Financial Services Firm (name withheld under NDA)
Results
- 70% Documents cleared without human review
- 34 → 9 min Mean review time per file
- 99.2% Citation accuracy on audit sample
- 4.1M Pages indexed
- 0 Documents leaving the client VPC
- 2.8x First-year ROI
Overview
A regional commercial lender with roughly $9B in assets runs a credit-file review step on every loan origination and annual renewal. An analyst opens a file — typically 40 to 600 pages of tax returns, financial statements, entity documents, appraisals and correspondence — and answers a fixed checklist of 31 questions. Does the guarantor structure match the approved terms? Is the most recent financial statement within policy age? Are all required entity authorisations present and current?
Mean handling time was 34 minutes per file. The queue ran four to nine days behind through most of the year and considerably worse at quarter end. Hiring more analysts had been tried; the bottleneck was not headcount, it was that the work is slow, linear reading.
The problem
This is the use case every vendor demos and very few ship, because the demo and the production system have almost nothing in common.
Wrong answers are worse than no answers. An analyst who misses an expired authorisation creates a remediable exception. A system that confidently asserts an authorisation is present when it is not creates a regulatory finding. The system had to be able to say “not found in this file” and be trusted when it said so.
Every answer needed a citation an examiner could follow. Not a similarity score. A document name, page number and bounding region that a human can open and verify in seconds. Without that, the output is unusable in a regulated review no matter how accurate it is.
A third of the corpus was scanned paper. Including faxes of faxes, handwritten annotations in margins, and tax returns photographed at an angle. Generic text extraction silently dropped content from these — the worst possible failure mode, because the pipeline reports success.
Nothing could leave the premises. The client’s third-party risk policy and their regulator’s expectations ruled out hosted model APIs for documents containing customer financial data. That made open-weight models served on client hardware the only viable path, decided in week two rather than discovered in week eight.
Approach
Week 1 — Discover. We shadowed six analysts through 40 files and did something we now do in every retrieval engagement: we built the gold set before building the system. Three senior analysts answered all 31 checklist questions for 120 representative files, with citations, independently. Their disagreement rate with each other was 7% — which told us precisely how much of the remaining error budget was irreducible ambiguity in the checklist itself rather than model failure. Six of the 31 questions turned out to be ambiguously worded; fixing the wording was free accuracy.
Week 2 — Architect. We split the 31 questions into three tiers by what they actually require: deterministic lookups (document presence, date arithmetic) that should never touch a language model; single-passage extractions; and genuinely synthetic questions requiring reasoning across several documents. Roughly 40% of the checklist fell into tier one. Routing those to ordinary code rather than a model is the least glamorous and most valuable architectural decision in the project.
Weeks 3–4 — Build. The ingestion pipeline does the heavy lifting. Layout-aware parsing preserves table structure and reading order; PaddleOCR handles scanned pages with a confidence gate that routes low-confidence pages to a human queue rather than passing degraded text downstream silently. Chunking follows document structure — a financial statement table is one chunk, not four fragments split at an arbitrary token count. Retrieval is hybrid: dense vectors catch paraphrase, BM25 catches exact entity names and account numbers, fused with reciprocal-rank fusion and reranked. Dense-only retrieval, which we benchmarked, missed exact-match entity lookups badly enough to disqualify it.
Generation runs under constrained decoding against a Pydantic schema per question type. Every response must include at least one citation resolvable to an indexed chunk; a response that cites nothing is rejected by the validator and retried, then escalated to a human. There is no path by which an uncited claim reaches an analyst.
Week 5 — Validate. Evaluation against the week-one gold set, question by question, measuring answer accuracy, citation accuracy and abstention behaviour separately. We deliberately over-weighted abstention: a system that abstains on 20% of questions and is right on the rest is far more valuable here than one that answers everything at 90%. Red-teaming included adversarial files — expired documents that look current, near-duplicate entity names, a renewal file containing a prior year’s statement as a decoy.
Week 6 — Deploy. Live in shadow mode across the full analyst team, then progressive enablement by question tier over five weeks as confidence accumulated per tier.
Solution
The analyst now opens a file to a pre-populated checklist. Each answered question carries its citation as a clickable link that opens the source page with the relevant region highlighted. Questions the system abstained on are flagged at the top of the queue with the retrieved candidate passages attached — so even an abstention saves reading time.
Three properties make it defensible in an examination:
Citation enforcement at the schema level. Not a prompt instruction, a validator. Uncited output cannot reach the UI.
Calibrated abstention. Per-question-type confidence thresholds tuned on the gold set, with abstention routing to the human queue. The system is designed to be wrong rarely rather than helpful often.
A complete audit trail. Every answer records retrieved chunk ids, reranker scores, model version, prompt version and timestamp to an immutable store. When an examiner asks why the system said what it said in March, there is an answer.
Results
Measured over the first two quarters of full production use:
- 70% of checklist questions cleared without human review, concentrated in tiers one and two.
- Mean review time from 34 minutes to 9, a 74% reduction in analyst handling time per file.
- Citation accuracy of 99.2% on a 500-question quarterly audit sample, independently checked by internal audit.
- Answer accuracy of 96.4% on non-abstained questions, against a 93% human-consistency baseline.
- Abstention rate of 11%, all routed to analysts with retrieved context attached.
- Queue backlog from 4–9 days to same-day through quarter end.
- Zero documents left the client VPC.
- 2.8x first-year ROI, with the renewal-cycle capacity increase as the dominant term.
Internal audit’s review of the system was completed without findings — which the Head of Credit Operations described as the outcome that mattered most.
Technology stack
Fully self-hosted inside the client’s VPC on Kubernetes. Open-weight generation model served with vLLM on client-owned GPUs. Source code, model weights, evaluation gold set, Terraform modules and the audit-log schema were delivered to the client at close, with no runtime dependency on KODA.