Financial Reporting Council Publishes Generative and Agentic AI Guidance: Risks, Mitigations and Illustrative Examples
The UK's audit regulator has quietly published a manual for how firms should use AI to check company accounts and it reveals just how much human oversight is still required to make the technology safe enough to trust.
The Financial Reporting Council's March 2026 guidance on generative and agentic AI in audit sets out three ways AI tools can fail in this context (p.6):
- The output itself can be wrong, through hallucination, omission, distortion, faulty reasoning, or internal inconsistency (p.8)
- The output can be correct but misused, because the auditor misunderstands what it means or how it fits into the broader methodology (p.18)
- The firm's entire approach can be non-compliant with auditing standards, even when the tool performs exactly as designed (p.19)
The document is frank about the limitations of the underlying technology (p.8). Large language models generate text by recognising statistical patterns, not through genuine understanding. They can sound confident when they are wrong. They have finite context windows, meaning detail gets lost in long tasks.
In agentic systems, where AI orchestrates multi-step work without step-by-step human oversight (p.4), the guidance identifies a specific additional risk: the system's understanding of its own goal can drift as earlier context gets compressed. A flaw introduced at the goal-interpretation stage can quietly corrupt everything that follows (p.11).
The proposed solutions are methodical:
- Certify tools per use case, testing for edge cases, known biases, and diverse scenarios (p.32)
- Train staff to recognise automation bias and not to wave through AI outputs uncritically (p.35, p.38)
- Build human review into control points throughout agentic workflows (p.37-38)
- For third-party AI components whose internal logic firms cannot inspect, obtain vendor assurance, while acknowledging that even then, firms may want to run their own tests (p.31)
The document is explicit that tolerating some level of deficient AI output is unavoidable and that deciding how much deficiency is acceptable is "a matter of professional judgement" (p.21, p.32).
We'd like your views:
- The FRC places ultimate responsibility for AI output quality on the audit firm's "professional judgement" (p.21). If an AI-assisted audit misses a material misstatement, where does liability sit, with the firm, the AI vendor, or the methodology team that certified the tool?
- The guidance identifies goal drift in agentic systems as context gets compressed across long tasks (p.11). At what point does that drift become an audit failure rather than a system limitation. Who decides?
- "Human in the loop" is recommended throughout, but the guidance flags automation bias as a real risk (p.16, p.35, p.38). Is mandating human review meaningful if the reviewer is already inclined to trust the machine oir wave through unchecked content when under pressure to complete tasks?
- The FRC describes tolerance for deficient AI outputs as "a matter of professional judgement" (p.21, p.32). Should audit clients, whose financial statements depend on these tools, have the right to know what deficiency tolerance was applied and when AI was used?
- Where firms cannot evaluate the internal logic of third-party AI components and rely on vendor assurances instead (p.31), does that create a structural dependency that sits outside the current regulatory perimeter?
Source: Financial Reporting Council, Generative and Agentic AI Guidance: Risks, Mitigations and Illustrative Examples, March 2026 URL: https://aboutblaw.com/blju


