A retrieval-grounded assistant that reads 400-page credit files and drafts decisions with citations, lifting analyst throughput 3.4×.
Where it started
Credit analysts were spending most of a working day per file, manually reconciling financial statements, bank data and legal documents that arrived as scanned PDFs. Volume had tripled in eighteen months and hiring could not keep pace. Any automation had to be reconstructable: the regulator's standing question is why a specific decision was made on a specific day.
- 01
Built an ingestion pipeline that parses scanned statements, tables and legal annexes into structured, permission-tagged chunks with page-level provenance.
- 02
Implemented hybrid retrieval with a reranking stage, tuned against 1,200 analyst-labelled question–answer pairs rather than a generic benchmark.
- 03
Generated draft assessments where every claim carries a citation to the source span, and refusal is the default when retrieval confidence is low.
- 04
Wired an evaluation harness into CI so prompt or model changes cannot regress accuracy without failing the build.
- 05
Ran an adversarial pass for prompt injection through uploaded documents, then converted each successful attack into a permanent regression test.
- Claude
- Postgres + pgvector
- Python
- Next.js
- AWS
Where it landed
Analysts now review and correct a draft rather than assembling one. Median time per file fell from roughly seven hours to just over two, with decision quality measured as equal or better on a blind sample. The system passed SOC 2 Type II audit with the model change log accepted as evidence.
They rebuilt our inference stack in eleven weeks and cut cost per request by 62%. The handover documentation was better than ours.
A short conversation with an engineer, not a sales qualification call. If we're the wrong people for it, we'll say so and point you somewhere better.