FCA’s Mills Review should focus on AI failure modes, not “innovation governance”
The FCA’s Mills Review into AI in retail financial services should be read as a warning. The sector is moving from "should AI be used?" to "what happens when it goes wrong at scale?"
Most firms still treat model risk as paperwork: a policy, a review, a sign-off, then business as usual. That breaks down when AI is embedded into affordability checks, fraud flags, customer comms, complaints handling, and collections. Under pressure, systems fail through brittle data dependencies, silent drift, and automation that outruns human escalation.
The failure mode is rarely a headline-grabbing meltdown. It is a slow accumulation of small errors that push customers into worse outcomes, then get defended as "the model said so".
The real test is control, not cleverness. Can a firm trace a customer outcome back to model version, inputs, and rules? Can it pause a model without breaking the service? Can it spot when staff start deferring to outputs they do not understand?
A good review should force the industry to stop hiding behind principles-based comfort. Evidence beats intent. Logs beat slides. Rehearsed rollback beats post-incident apologies.
We'd like your views:
- Which retail finance uses should face the strongest constraints: credit, collections, fraud, or advice?
- What is the minimum evidence a firm should hold before an AI model influences customer outcomes?
- Should AI vendor contracts be required to support forensic access to prompts, routes, and versions?
- How should regulators test for automation bias inside frontline teams?
- What does a realistic kill switch look like when models are embedded across multiple systems?

