Written by a human
Explainable AI in surveillance: Model governance, testing, and the new "prove it works" standard
The FCA received 3,806 STORs in 2025 and found real weaknesses in how firms govern the models meant to catch that risk. This guide covers what explainable AI in surveillance means, the governance lifecycle regulators expect, and how firms evidence it.
Regulators have changed the question. It’s no longer enough to have a surveillance system. Firms are now expected to prove it works: that their models are explainable, tested, calibrated, and governed across their lifecycle. Explainable AI in surveillance is the ability to show how a model reached an alert, or why it missed one, and to evidence that the system detects the risks it’s meant to.
That shift isn’t rhetorical. In its March 2026 Wholesale Markets Regulatory Priorities report, the U.K.’s Financial Conduct Authority (FCA) told firms plainly that having effective systems and controls to identify, assess, and mitigate risk is what meeting legal and regulatory obligations now requires. Where it finds serious failings, the regulator will consider every tool available, including enforcement action.
Critically, the presence of a surveillance platform is no longer treated as evidence of control. What matters now is whether a firm can demonstrate, with evidence, that detection actually works across its data, its models, and its operations.
This article sets out why that shift happened, what explainable AI in surveillance mean, and the governance lifecycle regulators expect firms to run. Importantly, it will also outline the frameworks converging on that expectation, and what “prove it works” looks like in practice.
Key takeaways
- Regulators’ expectations extend well beyond simply having a surveillance system in place. Firms must be able to show their surveillance model governance is demonstrable and defensible.
- Incomplete or inaccurate data feeds, ineffective alert calibration, and weak testing and governance of surveillance models are some of the regulators’ key concerns.
- Explainability isn’t the same as accuracy. A surveillance model can perform well and still be indefensible if no one can explain or evidence how it decides.
What regulators now expect: Prove it works
Historically, supervisory attention on market abuse controls centered on whether a firm had deployed a surveillance system covering the right instruments and communication channels.
That bar has moved. Supervisors now assess surveillance the way they assess any other control: by whether it functions as intended, under examination, over time.
The FCA leads that shift, with parallels across the U.S. Securities and Exchange Commission (SEC), Financial Industry Regulatory Authority (FINRA), the Commodity Futures Trading Commission (CFTC), and the Bank of England. The practical question firms face is no longer whether they have a system, but how to prove their surveillance model works.
The scale of what surveillance is expected to catch makes the stakes concrete.
According to the FCA’s Regulatory Priorities: Wholesale Markets report, the regulator received 3,806 suspicious transaction and order reports (STORs) in 2025, with 82% attributed to insider dealing.
Against that volume, the FCA’s own supervisory work found real weaknesses: incomplete or inaccurate data feeds, ineffective alert calibration, and weak testing and governance of surveillance models.
What “explainable AI in surveillance” means
Definition: Explainable AI in surveillance is the ability to show how a model reached an alert (or why it didn’t) with traceable, reviewable records, so the decision can withstand examination by internal audit, senior management, and regulators.
Explainability is distinct from accuracy. A model can perform well on aggregate detection metrics and still be indefensible if no one in the firm can explain or evidence how it reached a specific decision. Regulators, internal audits and, ultimately, the model’s own users need to be able to trace a path all the way from input data to alert (or to a deliberate non-alert), not simply trust the output. That traceability is what separates a genuinely governed model from a black box that happens to work most of the time.
This matters more, not less, as surveillance shifts from static rules-based logic toward machine learning and large language models capable of reading full conversations for context. The more sophisticated the model, the more the burden of proof shifts onto the firm to show its reasoning. That’s why AI model validation, testing a model’s behavior before and after deployment rather than assuming it works, sits at the heart of the lifecycle.
The pillars of demonstrable surveillance
Regulatory expectations for surveillance model governance now map onto a recognizable lifecycle, running from before a model is deployed through to its retirement. Each stage, from surveillance model testing to ongoing calibration, produces evidence that a firm can point to when asked to prove the system works.
| Expectation | What it means in practice |
| Model inventory and classification | Know every surveillance model in use and its risk level. |
| Validation before deployment | Carry out independent surveillance model validation before launching and document it. |
| Calibration and threshold tuning | Set and revisit alert thresholds so genuine risk isn’t drowned out or missed. |
| Ongoing monitoring | Watch for surveillance model drift and degrading performance; re-test on a regular basis. |
| Explainable audit trail | Be able to show how a model reached, or missed, an alert, with traceable records. |
| Governance and oversight | Formal approval, human oversight, and board-level reporting on effectiveness. |
None of these stages works in isolation. A model well-calibrated at launch but never re-validated will drift, and a clean audit trail with no senior sign-off has no accountable owner. Firms that can evidence all six are the ones equipped to answer a supervisor directly.
The frameworks that shape it
No single rulebook defines explainable AI in surveillance. Instead, several frameworks, some binding and some voluntary, are converging on the same expectation for model risk management for surveillance. What these all have in common is the expectation that models are governed, tested, and explainable across their lifecycle.
Together they form the backbone of trade surveillance model governance, and firms need to track each one, because the frameworks are themselves moving targets.
| Framework | What it contributes |
| SR 11-7 / SR 26-2 (U.S. model risk management) | SR 11-7 was the baseline U.S. discipline for developing, validating, and monitoring models before the Federal Reserve, the Office of the Comptroller of the Currency, and Federal Deposit Insurance Corporation replaced it with SR 26-2 (OCC Bulletin 2026-13) in April 2026. This updated guidance carries forward the core requirements for sound development, documentation, validation, monitoring, and governance. However, it explicitly places generative and agentic AI outside its formal scope, though it states that existing risk-management principles should still apply to those tools. |
| EU AI Act | Sets transparency, logging, and human-oversight obligations for AI systems, including high-risk AI systems. Firms should track which obligations apply on which date rather than treating the Act as a single deadline. |
| FCA (AI principles; SYSC; SM&CR) | Principles-based expectations that AI-driven surveillance is governed, tested, and effective, with senior individual accountability under SM&CR, sitting alongside the detection obligations firms already have in place under the Market Abuse Regulation (MAR). The FCA’s AI Lab, including its AI Live Testing service, gives firms a route to test AI systems under regulatory engagement before wider deployment. |
| NIST AI Risk Management Framework; ISO/IEC 42001 | Voluntary frameworks that many firms use to structure governance ahead of, or in parallel with, binding requirements: the NIST AI RMF for identifying and managing AI risk, and ISO/IEC 42001 for running a certifiable AI management system. |
The calibration problem
Surveillance alert calibration sits at the center of the FCA’s concerns for a practical reason, since it’s where good surveillance design most often breaks down.
A model calibrated too loosely floods analysts with false positives. The problem with false positives is that they bury genuine risk in noise while training reviewers (consciously or not) to work through alerts faster and with less scrutiny. In contrast, a model calibrated too tightly quietly stops generating the alerts that matter.
Neither failure looks dramatic day to day, which is exactly why the FCA singled out ineffective alert calibration as a specific finding rather than a hypothetical risk. It’s a slow, structural failure that’s easy to miss until a STOR that should have been raised didn’t get raised.
Reducing false positives in trade surveillance is often framed purely as an efficiency play. Fewer alerts means less analyst time. But done properly, false positive reduction is a risk-control outcome, not just a productivity one.
Calibration is also not a one-time configuration step. Trading behavior, market structure, and evasion techniques all change over time, so thresholds appropriate a year ago may not be appropriate now. Regulators expect firms to test, tune, and evidence that tuning on an ongoing basis, rather than set thresholds once and assume they still hold.
How firms build demonstrable, explainable surveillance
Running the lifecycle above consistently is what turns surveillance model governance from a paper policy into something a firm can demonstrate on request. In practice, that means treating each stage as a source of evidence:
- A documented validation report before go-live;
- A record of every calibration change and the reasoning behind it;
- And an audit trail built into the alert itself rather than reconstructed after a regulator asks for it.
All of this needs a named accountable owner and board-level visibility.
This model fits perfectly with Global Relay’s surveillance solution, which was built for demonstrable, defensible detection. Specifically, it:
- Uses generative AI to explain the reasoning behind each alert through a visible chain-of-thought
- Evaluates risk as part of message classification (rather than as an after-the-fact justification)
- Produces audit-ready reporting so firms have evidence of tuning and effectiveness on hand when a supervisor asks for it
The next frontier
Looking to the future, there’s more to build into surveillance model governance. Agentic AI models take multi-step actions rather than simply flagging output for human review. Plus, they sit outside the scope of the current U.S. model risk guidance and ahead of most binding EU AI Act obligations. This means that firms deploying them are largely governing them against voluntary principles of their own design.
This challenge extends across channels. Voice compliance monitoring and text-based surveillance converge on the same AI models, and continuous model observability, rather than monitoring only at scheduled checkpoints, is becoming a practical necessity as models grow increasingly autonomous.
Importantly, regulators are building their own evidence base for what tested and governed AI looks like in practice. The FCA’s AI Live Testing program enables firms to test AI systems in real-world conditions, with appropriate regulatory support and oversight.
FAQs
What is explainable AI in surveillance?
Explainable AI in surveillance is the ability to show how a model reached an alert, or why it didn’t, with traceable records that can withstand scrutiny by internal audit, senior management, or a regulator.
What does “prove it works” mean for surveillance?
Proving a surveillance model works means demonstrating its effectiveness to regulators — that it is validated, calibrated, monitored, and evidenced — rather than simply pointing to the fact that a system has been deployed.
How do you validate a surveillance model? What about calibration?
Validation means independently testing a model before it goes live and documenting that testing. Calibration means setting and continuously revisiting alert thresholds so genuine risk isn’t missed or drowned in false positives, with the tuning evidenced over time.
Does SR 11-7 apply to surveillance models?
SR 11-7 governed U.S. model risk management for fifteen years and was replaced in April 2026 by SR 26-2 (OCC Bulletin 2026-13), which carries forward its core validation and governance discipline for traditional quantitative models while placing generative and agentic AI outside its formal scope.
What are the FCA’s explainability requirements for surveillance models?
FCA explainability requirements don’t align with a single rulebook. Instead, its principles-based expectations under SYSC and SM&CR require firms to show governed, tested, and effective surveillance with clear senior accountability. The FCA’s supervisory findings on weak model governance and testing set the practical bar for what that evidence should look like.
Because poorly calibrated models fail quietly. Too loose and they bury real risk in false-positive noise, but too tight and they miss it altogether. This is a specific, recurring weakness the FCA has identified in firms’ surveillance frameworks.
What is surveillance model drift and why does it matter?
Surveillance model drift is the gradual decline in a model’s accuracy as market behavior, trading patterns, or evasion techniques change after deployment. That’s why ongoing monitoring and re-testing, not just initial validation, are part of the expected governance lifecycle.
The bottom line
Deploying a surveillance system is no longer enough. Firms must be able to prove it works, and explainability, testing, and model governance applied consistently across the model lifecycle are how they do it. While the frameworks defining that bar will keep shifting, the underlying discipline of validating, calibrating, explaining, and governing surveillance models won’t.
To see how demonstrable, explainable detection works in practice, explore Global Relay’s wider AI compliance solutions or speak to a member of the team.