TL;DR: Transaction monitoring rule tuning is the process of adjusting scenario thresholds, parameters, and customer segments so that each rule generates alerts worth investigating without suppressing genuinely suspicious activity. Most institutions skip it or do it badly: Capgemini's KYC/AML benchmark study found an average transaction monitoring false positive ratio of 92 percent. Done properly, tuning combines above-the-line and below-the-line sampling, sensitivity analysis, and segmentation, and produces the documentation examiners expect under the revised interagency model risk guidance and FinCEN's proposed AML/CFT program rule.
What Rule Tuning Actually Changes
Rule tuning changes the settings of detection scenarios that already exist. It does not write new scenarios or decide what happens to an alert once it fires. A structuring scenario has a dollar threshold, a lookback window, a minimum transaction count, and often a customer segment it applies to. Tuning is the work of deciding what those values should be, testing the decision against real data, and documenting why the institution landed where it did.
Every vendor system ships with default parameters set to work tolerably across many institutions, which means they were not set to work well for any particular one. A $10,000 aggregate cash threshold behaves very differently at a community bank with a retail deposit base than at a payments company whose merchants routinely move six figures a day. The rule logic may be identical. The right threshold is not. For a grounding in the engine itself, Sphinx has covered how AML transaction monitoring works separately.
Why Untuned Rules Fail
Untuned rules fail in both directions at once. They alert on legitimate behavior because thresholds sit below the normal range for the customers they cover, and they miss suspicious behavior because thresholds sit above the range where laundering actually occurs for others. The false positive problem is visible in the queue. The false negative problem is invisible until an examiner or a lookback exposes it.
Capgemini's KYC/AML benchmark study reported an average false positive ratio of 92 percent, with some institutions approaching 100 percent, and found no correlation between the degree of tool customization and the false positive ratio. An Everest Group assessment published in December 2025 put false positives at 85 to 90 percent of total alerts and found that only 2 to 4 percent of alerts were flagged as actionable for escalation or SAR filing.
The customization finding matters: writing more rules does not fix calibration. Three root causes account for most of the damage. Thresholds set once at implementation and never revisited. A single threshold applied across populations with different transaction profiles. And no feedback loop from investigation outcomes back into the parameters.
Tuning Rules vs. Triaging Their Output
Tuning and triage solve different problems. Tuning operates upstream, on the rules, and determines how many alerts exist. Triage operates downstream, on the alerts, and determines how fast and how well each one gets reviewed. An institution with badly tuned rules and excellent triage clears its queue efficiently while still reviewing thousands of alerts that never needed to exist. An institution with well-tuned rules and no triage capacity generates a smaller queue and still falls behind on it.
Sphinx has written about what actually reduces false positives, and the honest answer is that rule calibration and alert-level automation address different layers of the same cost. Rule tuning also has a ceiling: no threshold can distinguish a legitimate $9,800 cash deposit from a structured one. That determination requires context about the customer and counterparty, which is investigative work rather than parameter work. Teams that try to solve a growing alert backlog purely by raising thresholds end up with a defensibility problem when someone asks what was in the alerts that stopped firing.
Above-the-Line, Below-the-Line, and Sensitivity Analysis

Above-the-line and below-the-line testing is the core evidentiary method in threshold tuning. The production threshold draws a line through the transaction population. Above-the-line testing raises the threshold and reviews a sample of the alerts that would be suppressed, to confirm no productive alerts sit in that band and to quantify how much false positive volume could be shed. Below-the-line testing lowers the threshold and reviews a sample of the additional alerts that would be created, to determine whether the production setting is missing activity the institution should be seeing.
Neither test substitutes for the other. ATL alone tells the institution it can raise a threshold safely. BTL alone tells it whether it should have lowered one. Running both converts a threshold change from an assertion into an evidenced decision. Sample sizes should be set statistically, sampled alerts dispositioned by qualified analysts using production standards, and results recorded per alert so the yield of each band can be calculated.
Sensitivity analysis is the complement. Where ATL/BTL examines specific alternative thresholds, sensitivity analysis maps how alert volume and yield respond across a range of parameter values, and how parameters interact. A scenario whose output swings wildly with a small change in one input is unstable, and that instability is itself a finding. The interagency model risk guidance treats sensitivity analysis as a standard validation technique.
Segmentation Decides Whether Thresholds Mean Anything
A threshold is only meaningful relative to a population. Segmentation divides customers into groups with comparable transaction behavior so that each group gets thresholds calibrated to its own normal range. Without it, the institution must choose between a threshold low enough to catch unusual retail activity, which drowns the queue in corporate false positives, and one high enough to spare corporate customers, which lets retail anomalies through.
The Hong Kong Monetary Authority's 2024 thematic review of transaction monitoring systems found that inappropriate segmentation led directly to less effective threshold setting, higher false positive volumes, and in some cases higher-risk activity being left unmonitored. Segmentation is a precondition for tuning to work, not a refinement layered on top of it.
Useful dimensions include customer type, product, expected activity from onboarding, geography, and risk rating. Segments should be large enough to support statistical threshold-setting and homogeneous enough that one threshold makes sense for everyone in them. Segments also drift: a customer onboarded as a small business that has grown into a mid-market payments operation is sitting in the wrong segment, generating alerts a re-segmentation would have eliminated.
How Often to Tune, and What Forces a Re-Tune
Periodic tuning should run on a defined cadence. The HKMA thematic review found that the institutions it examined ran periodic reviews at intervals of 12 to 24 months depending on size and complexity, and U.S. practitioner guidance clusters around 12 to 18 months. Scenario-level KPIs for alert volume, false positive rate, and SAR conversion should be monitored continuously between full reviews. Scheduled reviews are the floor. These events should trigger an out-of-cycle re-tune:
What Examiners Expect to See on Paper
Examiners expect a tuning decision to be traceable from the risk assessment through the analysis to the approved change and its post-implementation monitoring. The regulatory framing shifted in 2026. On April 17, 2026, the Federal Reserve, OCC, and FDIC issued SR 26-2, Revised Guidance on Model Risk Management, which supersedes both SR 11-7 and the April 2021 interagency statement on model risk management for BSA/AML systems. The revised guidance is risk-based and states that it is most relevant to banking organizations above $30 billion in total assets, but the principles of conceptual soundness, ongoing monitoring, outcomes analysis, effective challenge, and documentation carry forward.
The 2021 statement put the practical expectation plainly: prudent risk management of an automated transaction monitoring system involves periodically reviewing and testing the filtering criteria and thresholds to ensure they remain effective, and independently validating the system's methodology. That expectation now sits inside the revised guidance for institutions that treat their monitoring system as a model, and inside the FFIEC BSA/AML Examination Manual's independent testing pillar for everyone else.
FinCEN's proposed AML/CFT program rule, issued April 7, 2026, requires internal controls to be reasonably designed and kept current as the institution's risk profile changes, and states that a program that was once reasonably designed may cease to be so if it is not updated. A scenario running on a threshold set for a customer base that no longer exists is a concrete example. Sphinx has covered AML model validation requirements and the effectiveness-based program rule in more depth.
At minimum, the tuning file for each scenario should contain the risk assessment linkage, the segmentation logic, the statistical basis for current thresholds, the ATL and BTL sample results with per-alert dispositions, the sensitivity analysis, the approval record, and the post-change monitoring results. If a vendor performed the tuning, the institution still owns the documentation and should be able to explain it without the vendor in the room.
The Mistakes That Undo a Tuning Exercise
The most common failure is tuning to reduce volume without measuring what happened to yield. Raising a threshold will always cut alerts. Whether it cut productive alerts is the only question that matters, and an institution that cannot show the SAR and escalation yield of the suppressed band has not tuned anything. It has turned off detection and called it optimization.
Skipping below-the-line sampling is the second failure, common because BTL creates work and rarely produces good news. An institution that only runs ATL is only ever looking for permission to raise thresholds. Examiners notice when five years of tuning cycles have moved every threshold in the same direction.
Other patterns that undermine tuning: reviewing only the scenarios currently causing pain rather than the full in-scope library; treating alert-to-SAR ratio as the sole success metric when SAR decisions depend on investigative context; failing to retire scenarios that duplicate coverage and have never produced a productive alert; and tuning thresholds before validating that the data feeding the scenario is complete and correctly mapped.
How to Evaluate a Tuning Approach or Vendor
A credible tuning methodology, whether internal or purchased, can be tested against a short set of questions.
Any approach that promises a specific false positive reduction before looking at the institution's data should be treated with caution. The right reduction is whatever the yield analysis supports, and that number is unknowable in advance.
Where Sphinx Fits
Tuning determines how many alerts an institution generates. Sphinx operates on the alerts that remain. Sphinx's agents work inside the institution's existing monitoring and case management systems, review each alert against customer, counterparty, and transaction context, and document the reasoning behind every disposition in a form that supports audit and examination. That disposition data, recorded consistently at the alert level with the reason for each outcome, is the feedback loop most tuning programs lack. A well-tuned rule set and a well-documented triage layer are complementary, and institutions that invest in one usually find the return on the other improves.
Frequently Asked Questions
What is the difference between above-the-line and below-the-line testing?
Above-the-line testing raises a scenario threshold above its production setting and reviews a sample of the alerts that would be suppressed, to confirm none were productive and to quantify the false positive volume that could be removed. Below-the-line testing lowers the threshold and reviews a sample of the additional alerts that would be generated, to determine whether the production setting is missing suspicious activity. Both are needed for a threshold change to be evidenced rather than asserted.
How often should transaction monitoring rules be tuned?
Periodic tuning typically runs every 12 to 18 months for U.S. institutions, and the HKMA's 2024 thematic review found intervals of 12 to 24 months depending on size and complexity. Scenario-level KPIs should be monitored continuously between reviews, and new products, mergers, sustained volume shifts, new typologies, or examination findings should prompt an out-of-cycle review.
Does SR 11-7 still apply to transaction monitoring systems?
SR 11-7 was superseded on April 17, 2026 by SR 26-2, the revised interagency guidance on model risk management, which also replaced the April 2021 interagency statement on model risk management for BSA/AML systems. The revised guidance is most relevant to banking organizations above $30 billion in assets, but its principles still frame how examiners assess threshold tuning. Smaller institutions remain subject to the FFIEC independent testing expectation, which covers the effectiveness of automated monitoring.
Is a lower false positive rate always a sign of good tuning?
No. A false positive rate can be lowered by raising thresholds until almost nothing fires, which reduces noise and detection at the same time. Good tuning is demonstrated by yield, meaning the proportion of alerts that result in escalation or a SAR, measured in the bands that were added or suppressed.
Can rule tuning replace alert triage automation?
No. Tuning reduces the number of alerts that should never have fired, but no threshold can distinguish legitimate activity from suspicious activity that looks identical at the transaction level. That distinction requires investigative context about the customer and counterparty, which is triage work. The two address different layers of the alert cost.

.png)