Alexandre Berkovic and Chrisjan Wüst · Sphinx Frameworks · 2026
TL;DR: A published threshold is a published instruction. Rule-based monitoring still has a place, and it misses structuring, mule activity, and layering that never cross the line an adversary can read. Six behavioral lenses, read together against one customer, turn an alert into a narrative an examiner can follow.
Volume is not coverage
The $10,000 currency transaction reporting threshold has been public since 1970. So have the SAR thresholds that sit under it. When a control is a number the adversary can read, the adversary optimizes against it. FinCEN's SAR statistics for calendar 2025 show "transactions below the CTR threshold" as 8.65 percent of suspicious activity categories reported by depository institutions (FinCEN SAR Stats, via Forvis Mazars, 2026). The Bank Policy Institute's study of large U.S. banks found that 40 percent of SARs related to structuring, and that a median of 4 percent of SARs drew any follow-up from law enforcement (Bank Policy Institute, 2018).
Most alerts are noise. The BIS Innovation Hub's Project Hertha, run with the Bank of England, cites false-positive rates "as high as 95 percent" in bank alert systems (BIS Innovation Hub, 2025). A later BPI survey put the conversion at roughly 19 percent of AML alerts becoming cases, and the cost of a single SAR at an average of 21.41 hours (Bank Policy Institute, 2024). Set that against the filing volume: 4.7 million SARs and 20.5 million CTRs in fiscal 2024, about 12,870 SARs a day, with calendar 2025 setting a record at 4.105 million SARs (FinCEN, 2025). The Wolfsberg Group has said the rising volume of filings is not contributing proportionately to outcomes, and that an aim of "no SAR left behind" produces an ineffective system.
The largest monitoring penalties of the past three years were coverage failures. TD Bank's 2024 guilty plea and $1.8 billion criminal resolution, alongside a $1.3 billion FinCEN penalty, followed the exclusion of domestic ACH, most checks, and other transaction types from automated monitoring. The Department of Justice found that 92 percent of transaction volume, about $18.3 trillion, went unmonitored between January 2018 and April 2024, and that the bank added no new scenarios from 2014 through 2022 (U.S. Department of Justice, 2024; FinCEN, 2024). The FCA fined Metro Bank £16.7 million after a data-feed error left more than 60 million transactions, worth over £51 billion, unmonitored from 2016 to 2020. Its 2025 notice against Monzo found that because occupation, account purpose, and source of funds had not been collected, monitoring could not assess whether transactions were consistent with expected activity (FCA, 2025). Examiners ask whether the institution knows what normal looks like for each customer, whether every transaction type is seen, and whether the alert carries enough context to close. A threshold answers none of those.
Six lenses, one customer
FATF Recommendation 10 requires ongoing scrutiny of transactions to ensure they are consistent with the institution's knowledge of the customer, their business, and their risk profile. The U.S. SAR rule uses almost the same language: a transaction is reportable when it is not the type that particular customer would normally be expected to engage in. Both texts put the customer at the center. This framework organizes that reading into six lenses. They are inputs to one assessment, and they are not six independent alert queues.

Six lenses around one customer. Each produces context. The assessment sits at the center.
- Customer behavior: does this activity fit what this customer has done before, and what they said they would do?
- Transaction structure: is value being fragmented, aggregated, or passed through in a way that avoids a threshold or breaks an audit trail?
- Geographic factors: does the transaction touch a sanctioned or FATF-listed jurisdiction, directly or through a counterparty?
- Cash and fiat activity: is cash, conversion, or monetary-instrument activity consistent with the stated business?
- Counterparty profile: who is on the other side, and does their registry and digital footprint match the claim?
- Document fraud: where paperwork supports a transaction, is it usable, verifiable, and consistent with the wire?
A wire slightly above the customer's usual size means little on its own. The same wire, to a counterparty whose domain was registered last month, from an account dormant for six months, with an invoice whose value differs from the amount, means something. The alert that results carries a narrative: what was expected, what was observed, why the gap matters, and what has already been checked. Whether to file stays with the institution's investigators and its BSA officer or MLRO. Sphinx Frontline is the review layer that turns that narrative into a case a person can close.
What each lens is looking for
A baseline is a description of an account's own history, commonly built over a rolling six to twelve months and recalculated so it follows the customer. It covers value in and out, transaction count, the distribution of sizes, recurring counterparties, channels, and jurisdictions. For a business customer it sits beside the expected activity stated at onboarding. On day one there is no history, so the onboarding profile and a peer group carry the load, and the first months after onboarding are a high-attention window. Dormancy followed by movement is a question about what comes next, not a scenario name.
Structuring is the shape the threshold rules were written for, and it is also where they fail. Two deposits of $9,900 a few days apart should not be assumed to be structuring (FFIEC BSA/AML Manual, Appendix G). Transactions near $10,000 alone are not sufficient to require a SAR. The useful signal is fragmentation across accounts, branches, or conductors that resolves to one beneficiary, or value that passes through an account with no reason to be there. Smurfing is the same pattern with many hands. Layering is rapid movement that exists to break the trail.
Geography is two different questions. Direct exposure is a fact: the payment touches a comprehensively sanctioned program, or a listed party. Indirect exposure is a judgment about a counterparty's own footprint. After the June 2026 plenary, FATF's grey list stood at 22 jurisdictions, with Bosnia and Herzegovina and Iraq added and Algeria and Namibia removed. FATF does not call for enhanced due diligence on grey-listed jurisdictions, and it warns against de-risking (FATF, June 2026). Syria is on that grey list and is no longer under comprehensive OFAC sanctions. A program has to be able to hold both facts at once. The same discipline applies to cash, nested accounts, and trade documents: the declared business is the reference, and an invoice that disagrees with the wire is a finding, not a footnote. How those documents are tested is the subject of validating source of funds when the paperwork cannot be trusted.
From alert to a decision
The reviewer should see the profile, the deviation, which lenses contributed, and what has already been ruled out, before they start. Escalation runs from behavior to counterparty to source of funds. Low-risk alerts with complete, consistent evidence can be closed with an automated rationale that a human samples afterward. Disagreement between lenses, incomplete data, prior SAR history, or a potential filing decision goes to a person.
Sphinx argues the close through the same three roles used in screening: a Prosecutor for the case that the activity is suspicious, a Defender for the innocent explanation, and a Judge that arbitrates under the institution's policy. The pattern is described in the Interpretable Agentic Framework. It does not replace rules. Regulatory thresholds still have to be reported against, and some typologies are well described by a simple rule. Behavioral monitoring supplies the context that turns a threshold breach into a question with an answer, and it catches the patterns that never cross a line. Most institutions surveyed in 2025 expected to run a hybrid of rules and models. Name screening has the same requirement of an explained disposition, which is why this volume sits next to reducing false positives in watchlist screening and the broader argument in Financial Crime in the Age of AI.
Frequently asked questions
Does behavioral monitoring replace rules?
No. Thresholds still have to be reported against, and some typologies are well described by simple rules. Behavioral monitoring adds the customer context that turns a breach into a question, and it catches patterns that never cross a threshold.
How long does a baseline take for a new customer?
There is no baseline on day one. The onboarding profile and a peer group carry the load. Many programs treat the first three months as a high-attention period and consider a baseline stable after six to twelve months.
Are FATF grey-list countries high risk?
They are a risk factor. FATF does not call for enhanced due diligence on grey-listed jurisdictions and warns against de-risking entire classes of customers. Weight the factor in the institution's own assessment, and refresh the list after each plenary.
Read the handbook
Threshold calibration, cash and instrument activity, counterparty research, and the escalation path are in Behavioral Transaction Monitoring. Where the payment is a wallet rather than an account, the same six questions change shape. That volume is crypto transaction monitoring at block speed. Banks that want the alert written before an analyst opens it can see the workflow.


.png)