AI for Financial Services and Insurance
The sector where a model decides who gets access to money, and where that decision has to be explainable to the person refused, to a supervisor, and to a court. Fifteen verticals, written for the person who owns the model risk rather than the person who demos it.
Explainability is not a nice-to-have here. It is the product.
In most sectors a model that works and cannot be explained is an engineering preference. In lending it is unlawful. A creditor still has to give an applicant the specific principal reasons for an adverse action, and a model whose reasons cannot be stated does not discharge that duty however good its performance. The same logic runs through insurance pricing, through capital models, and through anything a supervisor can ask you to justify after the fact.
So the constraint runs backwards into the architecture. Feature choices, model class, monitoring, challenger models and documentation are all downstream of a question most teams ask far too late: when this refuses someone, what exactly will we tell them, and can we prove it was true? Answer that in week one and the rest of the programme is ordinary engineering. Answer it in month nine and you rebuild.
Where AI actually earns its place in financial services
Ranked by evidence rather than by pitch frequency. The maturity column is our own read and we will defend it; the catch column is the part the vendor deck leaves out.
| Where | What it does | Maturity | The catch |
|---|---|---|---|
| Fraud detection and prevention | Scores transactions and applications in real time against behavioural and network signals. | Proven | The clearest win in the sector, and notably carved out of the EU AI Act's high-risk credit category where fraud prevention is the primary purpose. |
| Document and process automation | Extracts and validates data from statements, contracts, claims, KYC packs and submissions. | Proven | Exception rate rather than accuracy rate determines the saving. Ninety-five per cent accurate still means a human touches one in twenty. |
| Customer service and servicing | Handles routine enquiries, explains products, routes complex cases. | Proven | Safe for servicing and dangerous near advice. The line between explaining a product and recommending it is thinner than product teams assume. |
| Credit risk and underwriting | Scores applications and prices risk on richer data than a traditional scorecard. | Strong | Works well and carries the heaviest regulatory load anywhere here — Annex III high-risk, adverse action reasons, disparate impact. |
| Claims handling | Triages, routes, estimates and detects claims fraud. | Strong | Triage and estimation are safe. Automated denial is where insurers get into trouble, and increasingly where legislators are looking. |
| AML transaction monitoring | Reduces alert volume and prioritises what investigators see. | Mixed | Over half of banks run false positive rates above twenty per cent, yet published evidence of measured improvement from machine learning is thin. Most claims are expectations. |
| Trading and execution | Signal generation, execution optimisation, market making. | Mixed | Genuinely effective for execution and microstructure. Alpha claims are unfalsifiable from outside and usually should be treated as such. |
| Autonomous financial advice | Systems that recommend rather than inform. | Not ready | A regulated activity with suitability duties attached. We build the research and the explanation, not the recommendation. |
The number that should temper every AML pitch you hear
Roughly half of banks report false positive rates above twenty per cent and a quarter above forty, with more than half spending an hour or more investigating each alert. That is a real and expensive problem. What is missing from the market is published evidence that machine learning has measurably fixed it — surveys capture what institutions expect AI to deliver rather than what it has delivered. Treat any vendor's alert-reduction figure as a claim to be tested on your own alerts in shadow mode before it becomes a business case. See fraud detection.
Six regimes, and one of them changed this year
Financial services AI sits under prudential model risk rules, consumer protection law, insurance conduct rules, data protection, operational resilience requirements and now general AI regulation. This is our reading as at September 2026 and we work alongside your risk, compliance and legal functions rather than in place of them.
Annex III, from 2 December 2027
AI evaluating the creditworthiness of natural persons or establishing credit scores is Annex III high-risk, as is AI used for risk assessment and pricing in life and health insurance. Fraud detection is carved out where fraud prevention is the primary purpose. Deployers in banking must also complete a fundamental rights impact assessment, and fine-tuning a vendor model on your own data can reclassify you from deployer to provider.
SR 11-7 replaced in April 2026
The joint OCC, Federal Reserve and FDIC Revised Guidance on Model Risk Management took effect on 17 April 2026, superseding SR 11-7 and OCC 2011-12. It shifts from prescriptive validation cycles to principles and materiality-based tiering, with lifecycle lineage, continuous drift monitoring, and effective challenge that is versioned and reproducible rather than a one-off memo. Generative and agentic AI are formally out of scope pending a separate consultation — supervisors are applying the principles by analogy in the meantime.
The guidance moved; the statute did not
The CFPB rescinded its AI adverse-action circulars in May 2025, which removed interpretive guidance and changed nothing statutory. ECOA and Regulation B still require the specific principal reasons for an adverse action, and a model whose reasons cannot be stated does not satisfy that. Fair Housing Act disparate impact exposure survives, and enforcement now runs through state regulators, HUD and private litigation.
Twenty-five states, four frameworks, one testing rule
Twenty-five states and the District of Columbia have adopted the NAIC model bulletin on insurers' use of AI, with California, Colorado, New York and Texas operating their own frameworks — and adoptions are not uniform, with Virginia requiring risk to be eliminated rather than mitigated and Connecticut adding annual certification. Colorado goes furthest: quantitative disparate-impact testing of facially neutral models, extended to private passenger auto and health, with full compliance required from 1 July 2026.
DORA, applying since January 2025
ICT risk management, incident classification and reporting, resilience testing and a register of third-party ICT arrangements. For AI this bites hardest on model and inference dependencies: a hosted model provider is a third-party ICT provider, with the concentration risk, exit planning and contractual requirements that follow.
Article 22, and the SCHUFA problem
Automated decision-making with legal or similarly significant effect carries specific rights under GDPR, and European case law has treated credit scoring as automated decision-making even where a human formally signs the outcome. A rubber-stamp review does not convert an automated decision into a human one, which is a design requirement rather than a policy statement.
What this means for a build
Three things are architectural rather than procedural, and retrofitting any of them is a rebuild. First, reason codes: every adverse decision must be able to state its specific principal reasons truthfully, which constrains model class and feature engineering from day one. Second, lineage: the new model risk guidance expects an unbroken chain from data through development to monitoring, versioned and reproducible. Third, meaningful human review: if a person's involvement is a rubber stamp, the decision is still automated. See EU AI Act compliance readiness and explainable AI.
Who we write for
Each page starts from that organisation's own problems, names the regulatory exposure it carries, and routes into the engineering. Depth varies and is stated on each page.
Retail banking
Fraud, credit decisioning, servicing and AML — with reason codes and model risk evidence built alongside the model.
Read more →Commercial banking
Credit memos, spreading, covenant monitoring and early warning — document work, not scorecards.
Read more →Credit unions
Member servicing, fraud and collections — where the core provider's module usually wins.
Read more →Insurance carriers
Claims, underwriting support and fraud — with disparate impact testing designed in, not bolted on.
Read more →Insurance brokers
Submissions, market matching, policy comparison and renewals — document work, with the advice line drawn.
Read more →Insurtech
Instant quote, embedded distribution and claims automation — plus the governance your carrier partner will demand.
Read more →Capital markets & trading
Execution, microstructure, surveillance and research automation — with alpha claims treated sceptically.
Read more →Asset & wealth management
Research, reporting, onboarding and adviser support — with the suitability line kept intact.
Read more →Private equity & VC
Sourcing, diligence document review, portfolio monitoring — and independent AI diligence on targets.
Read more →Fintech
Fraud, onboarding, credit decisioning — plus the compliance evidence your banking partner will require.
Read more →Lending & mortgages
Underwriting, document automation and servicing — with reason codes and disparate impact testing as design constraints.
Read more →Payments
Real-time fraud, authorisation optimisation, merchant risk and disputes — engineering under a latency budget.
Read more →Audit & assurance
Full-population testing, journal anomalies and document review — with the audit judgement staying human.
Read more →RegTech
Screening, surveillance, regulatory change monitoring — plus the validation evidence your customers need.
Read more →Cryptocurrency & digital assets
Chain analytics, transaction monitoring and surveillance — built for the compliance regime that has now arrived.
Read more →FAQ
Marked up with FAQPage schema so these answers can surface in search results and inside AI assistant responses.
Do you have financial services experience?
Yes in credit risk, fraud, document automation, claims and capital markets data engineering. Less in actuarial pricing, and none as a regulated entity ourselves — we build and document models that your model risk function validates and your regulator supervises, and we do not stand between you and either. Each of the fifteen vertical pages states our depth in that specific area rather than implying uniform expertise.
Will a machine learning model pass model validation?
It can, and whether it does depends far more on documentation and effective challenge than on architecture. The revised interagency guidance effective April 2026 is principles-based and materiality-tiered, which gives more room for non-linear models than SR 11-7 was often read to allow — provided you can evidence lineage, monitor drift against defined thresholds, and produce reproducible challenger comparisons. We build the evidence alongside the model, because assembling it afterwards is how validation cycles slip two quarters.
Can we use a black-box model for lending decisions?
Not for decisions you must explain, which in consumer lending is all of them. ECOA and Regulation B require the specific principal reasons for an adverse action, and post-hoc attribution methods produce explanations of the model rather than statements of fact about the applicant — a distinction that matters when the explanation is challenged. The workable pattern is a model class that yields honest reasons, or a constrained architecture where the reasons are structural rather than reconstructed.
How do we handle bias testing?
As a standing control rather than a launch gate. Disparate impact is measured on outcomes, not intentions, so a facially neutral model built on facially neutral data can still produce it — which is precisely what Colorado now requires insurers to test for quantitatively. We report performance and approval rates by protected class and proxy from the first evaluation, monitor them after deployment, and treat the feedback loop as a first-class risk. See AI bias audit.
What is the most common way financial services AI projects fail?
They build something accurate that cannot be put into production because model risk, compliance or legal were not in the room at the start. The second most common is a model that works in backtest and degrades on deployment because the training window covered a regime that has since ended. Both are avoidable, and both are cheaper to avoid in week one than to discover in month nine.
Start with the problem, not the technology.
Thirty minutes, no charge, no deck. Tell us what is going wrong in your organisation and we will tell you whether AI is the right instrument — including when it plainly is not.