Skip to content

AI in finance

Finance AI that reaches production and stays accountable after launch.

Most finance AI stalls after the demo. We pick the places where a model earns its keep, ground it in your own records, version every prompt and dataset, and watch drift once traffic is live. Reviewers keep a path to override, and exceptions route to a named owner.

Where it fits

  • The problem

    A pilot scored well on clean data, then met a quarter of messy statements and no one could say why an output changed. Meanwhile analysts still re-key figures out of statements by hand. The model sits in production with no paper trail, no owner, no record of which version made which decision.

  • What we build

    We build the retrieval layer over your policies and ledgers, the evaluation set that has to pass before release, and the monitoring that flags drift in inputs and outputs. Model versions, prompts and thresholds live in your repository. Every decision writes a record you can hand to an auditor.

  • How we engage

    Senior engineers do the work, no layered account team between you and the people writing code. The stack runs in your cloud or your data centre, with open weight models or a vendor model behind your own gateway. You keep the repository, the evaluation set and the runbook.

What we build

Six pieces of work that move a finance model out of the sandbox and keep it honest afterwards.

Use case triage

We rank candidate workflows by the cost of the current manual step and the tolerance for error. Reconciliation exceptions and policy lookups usually clear the bar. Credit decisions and client-facing advice get narrower scope with a reviewer in the loop, and some ideas get dropped before anyone writes code.

Grounding layer

A model with no source material guesses. We index your policy documents, contracts and ledger extracts, chunk them so a paragraph keeps its context, and pass retrieved passages with the query. Answers carry the passage and document version they came from, so a reviewer can check the source in one click.

Evaluation harness

Before release, the system runs against a labelled set drawn from your own awkward cases, including the statements that broke the last pilot. Accuracy, refusal rate and cost per task all get thresholds. A regression on any of them blocks the deploy and names the examples that failed.

Model governance

Each release records the model version, the prompt, the retrieval index snapshot and who approved it. Change one and the record forks. When a controller asks why a figure differed in March, the answer comes out of the log rather than someone's memory.

Drift monitoring

Input distributions move. New statement formats arrive, a counterparty renames its fields, and quality slips weeks before anyone files a ticket. We track input shape, confidence spread and reviewer override rate per workflow, and alert on the trend rather than the single bad day.

Exception routing

Low confidence output stops and waits. Items below the threshold queue for a named reviewer with the source passages attached, and the decision feeds back into the labelled set. Anything the model has no grounding for gets refused and escalated instead of filled in with a plausible number.

In production

What a governed finance model looks like

The demos below use synthetic portfolios and statements, with the same retrieval, thresholds and review queue we build for production.

How we work

  • Discover

    Systems, constraints, and the regulation you operate under get mapped before implementation starts, so the design accounts for what already runs and for what examiners will ask about.

  • Architect

    A design that fits your stack. Integration-first, self-hostable, and built to change as rules, regulation, and volume do.

  • Build

    Senior engineers ship in tight increments, each tested and reviewed as it goes, so the system is reviewable at every step instead of only at the end.

  • Harden

    Security, compliance, and load-testing run inside the build, so controls, audit trails, and peak-volume behavior are proven before launch.

  • Run

    The handover includes clean, documented systems, with the option to keep the same engineers operating them once they are live.

Why Oxagile

  • We know the finance domain

    Our engineers have shipped reconciliation, credit and payments systems, so the conversation starts with the control that has to hold rather than the model that looks interesting. We ask who signs off, what the audit expects, and which failure the business can absorb.

  • Governance built in from day one

    We build versioning, evaluation sets and monitoring alongside the first working prototype. Retrofitting a paper trail onto a live model costs more than building it early, and it rarely satisfies the people who have to sign the control narrative. Your auditors ask for both.

  • Yours to run afterwards

    Code, prompts, evaluation data and deployment scripts land in your repository under a licence your legal team has read. Your team can retrain, swap the model or turn the feature off without calling us. We document the runbook and the known failure modes.

20+
Years in software engineering
300+
Engineers
50+
Clients incl. Fortune 500

Questions

How do you stop the model making things up?

We control hallucination risk rather than claim to remove it. The model answers only from retrieved passages, cites the document and version behind each figure, and refuses when retrieval returns nothing relevant. Confidence thresholds send weak answers to a reviewer. The evaluation set includes questions with no valid answer, so refusal behaviour gets measured too.

Does our data go to a model vendor?

Only if you decide it should. We can run open weight models on your own hardware, or route to a vendor API through a gateway you control that logs every call and strips identifiers first. Residency, retention and the list of permitted endpoints get written into config, reviewed by your security team before the first production call.

What does the first phase look like and how long does it take?

A short discovery on one workflow, then a working slice against your own data. We agree the evaluation set and the pass thresholds before building, so the review at the end is a measurement rather than an opinion. Most first slices run a couple of months, and data access usually sets the pace.

Who owns the model risk once it is live?

Your risk function does, and we build the artefacts it needs. Each workflow gets an owner, a documented purpose, recorded limitations, monitoring thresholds and an off switch. Reviewer overrides and refusals are logged as evidence. We sit with your second line during the first review cycles so the control narrative comes from people who understand the system.

Part of AI in Finance

AI put to work in finance operations, grounded, governed, and running in production.

Get AI into production,
not just a demo.

Have a workflow begging for AI, whether documents, copilots or forecasting? Tell us the use case and we'll take it to production, governed for risk.