TL;DR: AML rule backtesting replays a proposed detection rule against historical transactions to estimate what it would have alerted on; shadow mode then runs the rule on live traffic without creating cases, for comparison with the incumbent before anyone acts on it. Backtests fail quietly through label leakage and survivorship bias, which is why a rule that looks precise on history can flood a queue in production. U.S. institutions filed 4.8 million SARs in fiscal year 2025 according to FinCEN's Year in Review; every untested rule change adds to that volume.
Backtesting Answers One Question. Shadow Mode Answers Another.

A backtest tells us what a rule would have done. Shadow mode tells us what it will do.
Backtesting executes a candidate rule, or a changed threshold, against a fixed window of historical transactions and customer data, producing hypothetical alerts that can be compared with what actually happened: which customers were investigated, escalated, or reported. Shadow mode, also called champion-challenger or a parallel run, deploys the new rule alongside the incumbent on live data. The challenger's alerts are logged but never routed to analysts, and after a defined period the two alert populations are compared. Shadow mode catches what a backtest cannot: feeds that behave differently in real time, drifted customer behavior, and how many alerts land on a Monday morning. Sphinx has covered how transaction monitoring generates alerts elsewhere.
Building a Backtest That Does Not Lie to You
The hard part of backtesting is constructing a dataset where the result means something. Three choices decide it.
The data window comes first. It has to cover seasonal cycles, and it has to be split in time, not at random. Calibrating a threshold on January through September and evaluating it on October through December mimics deployment; sampling evaluation rows from across the whole year does not, because the threshold has already seen the future it is being tested on. The revised interagency model risk guidance lists out-of-time testing alongside out-of-sample testing as a core development activity.
Label leakage is the second problem. The obvious ground truth is the case management system: which alerts became cases and which cases became SARs. The leak happens when the rule under test uses inputs that exist only because an investigation happened. Risk ratings raised after a SAR, flags added at escalation, counterparty tags applied during a lookback: all encode the outcome the rule is supposed to predict. Every field has to be reconstructed as of the transaction date.
Survivorship bias is the third, and the one we see go unexamined most often. Historical labels exist only for activity the incumbent rules caught, so a backtest scored against SAR outcomes measures whether the new rule finds what the old rules found, never whether it finds what they missed. A challenger designed to widen coverage always looks worse on precision than it deserves. Below-the-line sampling of historical activity that never alerted, dispositioned by analysts as if it were live, is the only way to put labels where none exist. Our rule tuning piece covers above-the-line and below-the-line sampling in depth; in a backtest, the below-the-line sample is what makes the coverage estimate honest.
What to Measure on Both Sides of the Line
A backtest report showing only alert volume has not measured anything a compliance officer can sign. Four metrics together describe a rule: alert volume by customer segment, as a change against the incumbent; precision, the share of alerts that would have become cases under production standards; SAR conversion, with the caveat that filing decisions depend on context the rule cannot see; and typology coverage, which of the institution's assessed risks the rule actually catches.
The fourth is the one most often skipped. A structuring rule with excellent precision that misses every funnel account case has a coverage gap, invisible unless the test set is tagged by typology. We keep confirmed cases and constructed scenarios per typology and run every candidate against them before looking at volume.
The Rule That Passed Backtest and Failed in Shadow
The pattern we see most often: a payments team proposes a velocity rule that alerts when a customer receives more than a set number of inbound transfers from distinct counterparties within a short window. On a twelve-month backtest it fires on a few hundred accounts, half already the subject of a case, with precision far above the incumbent library, and is approved for shadow.
In shadow it generates several times the backtested volume in its first two weeks. The backtest ran on settled transactions from the data warehouse; production ran on the real-time feed, which included pending and later-reversed transfers, so counterparty counts were inflated by retries. And the institution had launched a payroll product for gig-economy platforms eight months into the historical window. Those customers legitimately receive many small transfers weekly; they barely registered in the backtest and were a third of the covered population in live traffic.
The quieter version: a rule passes backtest and shadow, is promoted, and six weeks later almost every alert it produces is a customer the incumbent rules already flagged, because the shadow report never asked how many challenger alerts were unique. Overlap with the incumbent is the first number we read now.
Running Champion-Challenger Without Breaking the Queue
Shadow mode works when the comparison is designed before the challenger is switched on. The challenger runs on the production feed, at production cadence, with alerts written to a table analysts do not see, for a full seasonal cycle, typically four to eight weeks. Analysts then disposition a statistically sized sample from three populations, challenger-only, champion-only, and both, which show what the challenger adds, what it would lose, and where the two agree. The HKMA's 2024 thematic review of transaction monitoring systems found that functional testing, which compares a scenario's simulated output with the deployed system's, surfaced data-mapping and configuration errors that dashboards had not. Shadow mode is that test on live data, and it should reconcile inputs as well as outputs.
One obligation does not pause during shadow. If a sampled challenger-only alert reveals activity that meets the SAR standard, the institution knows about it and the filing clock runs, so shadow programs need a documented path for escalating such findings into the live case process.
Documenting It for Model Risk Management
The regulatory frame changed on April 17, 2026, when the Federal Reserve, OCC, and FDIC issued SR 26-2, Revised Guidance on Model Risk Management, superseding SR 11-7 and the April 2021 interagency statement on model risk management for BSA/AML systems; the OCC's Bulletin 2026-13 rescinds OCC 2011-12 and 2021-19. The revised guidance is most relevant to banking organizations above $30 billion in assets and excludes deterministic rule-based processes with no statistical theory behind them from its definition of a model.
Three things follow. A deterministic threshold rule at a smaller institution may sit outside the guidance's scope, but it remains inside the FFIEC independent testing pillar, and examiners still expect pre-deployment test evidence. A scored or machine-learning detection layer is a model under the guidance, which names back-testing as a form of outcomes analysis and says validation generally occurs before first use. And the December 2018 interagency joint statement on innovation still stands: pilot programs that expose gaps in a BSA/AML program will not necessarily result in supervisory action.
The test file for each change should let a validator reconstruct the decision without the team in the room: risk assessment linkage, the backtest window and label construction, leakage and survivorship controls, typology results, shadow sample sizes and per-alert dispositions, the overlap analysis, and the approval record. Sphinx has written about AML model validation requirements and independent testing expectations, which frame what that file is checked against.
Sign-Off and Change Control
A rule goes live when a named owner accepts documented results against criteria written before shadow started: maximum volume change, minimum precision, minimum unique coverage, no loss of any typology. Writing them afterward turns validation into rationalization. Change control ties the deployed configuration to the test file and compares the first thirty days of live output with the shadow projection, with a rollback path defined in advance.
Testing AI Triage the Same Way
An agent that decides which alerts are escalated and which are closed has the same failure modes as a rule, and we hold it to the same test. Backtesting an agent means replaying historical alerts with data reconstructed as of the alert date and comparing its dispositions with the analyst dispositions on record. The leakage risk is higher here: agents read case notes, and notes written after escalation are the purest form of leaked outcome.
Shadow mode for triage means the agent dispositions live alerts in parallel with analysts, its verdicts are hidden, and disagreements are adjudicated by a senior reviewer. The metric that matters most is the false discount rate: the share of alerts the agent would have closed that a human escalated. The HKMA review describes an institution tracking exactly this KPI for its machine-learning alert scoring. The revised model risk guidance places generative and agentic AI outside its scope, so the burden of proof falls on the institution's own evidence; a quality assurance program sampling agent dispositions after go-live is the ongoing monitoring half of it.
Where Sphinx Fits
When Sphinx's agents run in shadow alongside analysts, every disposition carries the reasoning and data it relied on under the Interpretable Agentic Framework, so the disagreement review above is a read of a log rather than a reconstruction. The agents work inside the systems an institution already uses, so the comparison runs on the same alerts analysts see.
Frequently Asked Questions
What is AML rule backtesting?
AML rule backtesting runs a proposed or modified transaction monitoring rule against a fixed window of historical transaction and customer data to estimate the alerts it would have generated. Those alerts are compared with recorded case and SAR outcomes to estimate volume, precision, SAR conversion, and typology coverage before deployment.
What is shadow mode in transaction monitoring?
Shadow mode, also called champion-challenger or a parallel run, deploys a new rule or model on live data alongside the incumbent while keeping its alerts out of the analyst queue. The two alert populations are compared over a defined period, typically four to eight weeks, before anyone acts on the challenger.
Why do rules that pass backtesting fail in production?
The most common causes are label leakage, where the backtest used fields that exist only because an investigation already happened; survivorship bias, where historical labels cover only what the old rules caught; and differences between warehouse data and the real-time production feed. Customer behavior can also drift, especially after new product launches.
Does SR 11-7 still govern testing of AML monitoring rules?
No. SR 11-7 and the April 2021 interagency statement on BSA/AML model risk were superseded on April 17, 2026 by SR 26-2, and the OCC rescinded Bulletins 2011-12 and 2021-19 at the same time. The revised guidance is most relevant to institutions above $30 billion in assets; smaller institutions remain subject to FFIEC independent testing expectations.

.png)