AML Quality Assurance: How to Build a Program That Holds Up

How to build an AML QA program: sampling design, scorecards, error taxonomy, 2026 enforcement actions, and what changes when AI agents disposition alerts.
Alexandre Berkovic

TL;DR: An AML quality assurance program samples closed alert dispositions and investigations after the fact to measure whether the first line applied policy correctly. It is a second-line control and does not satisfy the independent testing requirement. In August 2026 FinCEN imposed a $125 million penalty on UBS Financial Services — the largest BSA penalty ever against a broker-dealer — in a matter where a quality control gap in report generation caused roughly two years of missed alerts. No regulator publishes a benchmark error rate or minimum sample size, which means the design is yours to defend.

QC, QA, and Independent Testing Are Three Different Things

Three-column comparison of quality control, quality assurance and independent testing by line of defense, timing and owner
Quality control, quality assurance, and independent testing sit in different lines of defense and answer to different owners.

These get used interchangeably, and examiners notice when an institution cannot distinguish them.

Quality control is in-process and pre-decision. It is a maker-checker arrangement owned by operations, often applied to every item in a queue or targeted at open items before they close. It sits in the first line.

Quality assurance is post-decision and sampled. It reviews closed items to measure whether the first line applied policy correctly, and it is owned by compliance in the second line.

Independent testing is the fourth pillar. It assesses whether the program as a whole is adequate — including whether QA itself is working — and it reports to the board or a board committee comprised primarily or entirely of outside directors.

The FFIEC's independence constraint has a consequence institutions frequently get wrong in both directions. Persons conducting independent testing must not be involved in other BSA functions that present a conflict, and the manual names training and developing policies and procedures as disqualifying. So a QA team that also writes the procedures cannot be the independent testing function. Equally, having internal audit run QA compromises audit's independence over that same function. Presenting QA results as if they discharge the independent testing obligation is a straightforward finding.

On frequency, no statute specifies an interval. The FFIEC's stated sound practice is independent testing generally every 12 to 18 months, commensurate with the institution's risk profile — and more frequently after significant changes in systems, staff, or products. Institutions preparing for examination should be reading this alongside the rest of their BSA exam preparation.

Sample Closures, Not Escalations

Closed alerts feeding three separate sampling strata labeled risk-weighted, random and targeted
A defensible QA plan draws from closed alerts through three strata that are sized and reported separately.

The single most consequential design decision in a QA program is which population you draw from. Escalated alerts already get scrutiny — they go to a second-level investigator, and often to a SAR decision. The errors that matter are false negatives, and those live entirely in the closed population.

A defensible sampling plan runs three strata, and reports them separately:

Blending those into a single headline error rate is statistically meaningless, and an examiner with a quantitative background will say so.

On sizing, percentage-based sampling is the most common mistake. Precision in attribute sampling depends on absolute sample size, not on the share of the population. Reviewing 5 percent of 200 alerts is underpowered; reviewing 5 percent of 500,000 is wasteful. The useful heuristic is the rule of three: observing zero errors in a sample of n supports an upper 95 percent confidence bound of roughly 3/n on the true error rate. Zero errors in 30 items only establishes that the error rate is probably below 10 percent, which is not a reassuring statement to put in front of a board.

Worth saying plainly, because vendor content muddies it: no U.S. regulator publishes a benchmark AML QA error rate or a minimum sample size. The FFIEC makes sample size risk-based and discretionary. Any "industry standard 5 percent" figure is convention, not regulation, and framing it accurately is a credibility advantage on exam.

What to Score

Holistic pass/fail scoring drifts into meaninglessness within a couple of quarters. Score discrete dimensions that can be assessed independently:

Weight a decision that was correct but poorly documented differently from a decision that was wrong. Conflating them is the fastest way to make QA scores stop meaning anything.

The error taxonomy should separate findings by how they get remediated: critical findings where a SAR should have been filed and was not, substantive findings where the decision was defensible but evidence gathering was incomplete, documentation findings where the answer was right but the record is unreconstructable, process and SLA findings, and systemic findings where the defect is in the rule, threshold, data feed, or procedure rather than in the analyst.

That last category is not academic, and it is where the largest penalty in this space landed.

The $125 Million Spreadsheet

FinCEN's August 2026 consent order against UBS Financial Services is the clearest available illustration of why QA has to cover the data layer and not just the human decision layer.

An Excel-based report used to generate monitoring output systematically undercounted the value of foreign currency wires. The order states that "due to the lack of quality control in generating this report, UBSFS personnel failed to ascertain that UBSFS was systematically undercounting the value of these foreign currency wires for roughly two years," resulting in a failure to alert on hundreds of transactions. Staff who discovered the issue in late 2020 took no immediate steps to escalate it or assess the scope of the miss. A lookback happened only after a FINRA examination inquiry forced one.

The broader failure was larger: more than 61,500 foreign currency wires with an aggregate value above $10.5 billion went unmonitored for over four years — following a 2018 FinCEN consent order covering substantially the same defect.

No amount of analyst-level QA would have caught the Excel error. The population being sampled was already wrong before any analyst saw it.

The order also names investigation quality failures directly. In one case involving a Venezuelan financial institution, the investigator failed to obtain information identifying the originator of certain wire transfers, despite the missing originator being the primary reason the alert was escalated in the first place. In another, an investigator conducted no meaningful analysis of pass-through activity despite pass-through concerns being the basis for escalation.

The remedy is where FinCEN names QA as a required control, and it is worth quoting because it defines the standard: the order requires assessment of the processes second-level investigators use to review concerns raised by first-level reviewers — asking, for example, whether an alert escalated for potential pass-through activity produces an investigation that actually addresses pass-through activity — "as well as quality assurance, quality control, and/or independent testing to evaluate the consistency and effectiveness of such processes."

Two More Actions Worth Knowing

FinCEN's March 2026 order against Canaccord Genuity carried an $80 million penalty for at least 160 unfiled SARs. The mechanism is instructive: reviewers self-invented filtering thresholds. One employee, in their first couple of weeks on the desk and without supervisor guidance, used trial and error to arrive at what the order calls arbitrary numbers, filtering a 50,000-line daily report to manage the scope of the review without regard for whether the firm was still reviewing the activity the report was designed to detect. A daily wash sales report exceeding 50 pages was described internally as simply too long to review. Two compliance employees later falsified nearly 400 documents during an examination to create the impression reviews had been performed.

The lesson is not about bad actors. It is that with no second-line check, a first-line reviewer silently redefined the scope of an entire control, and nobody knew for years.

The OCC's April 2026 consent order against Community Federal Savings Bank is the most directly on-point action for automated dispositioning. The bank's automated alert triage system had deficiencies in its logic, data, and methodology that resulted in the system auto-closing alerts that should have been escalated — and, in the OCC's words, auto-closing "a very high percentage of all ingested alerts." The OCC separately found independent testing weak, with the internal auditor failing to identify program weaknesses and failing to scope high-risk areas into its work. This was a cease-and-desist with no monetary penalty, but it required a third-party SAR lookback consultant reviewing the quality and accuracy of previous filings.

The Model Risk Framework Just Moved

Anyone building QA over automated decisions should know that the ground shifted in April 2026. SR 26-2 superseded SR 11-7 and also rescinded SR 21-8, the BSA/AML-specific model risk statement institutions had relied on since 2021. The AML-specific model risk guidance no longer exists as standalone guidance.

More consequentially, SR 26-2 footnote 3 states that generative AI and agentic AI models "are novel and rapidly evolving. As such, they are not within the scope of this guidance." The revised definition also excludes deterministic rule-based processes, which arguably removes many rules-based transaction monitoring configurations from formal model risk management as well. The guidance is explicitly non-enforceable and aimed primarily at institutions above $30 billion in assets.

None of that touches the underlying legal obligation. 12 CFR 21.21, 31 CFR 1020.320, and the SAR filing duty apply regardless of what technology produced the decision. SR 26-2 says as much, directing institutions to let their own risk management practices determine governance for tools not covered.

The practical consequence is that QA becomes the primary defensible control over AI dispositions rather than a secondary one. With no validation framework attaching automatically, an institution's own sampling, scoring, and outcomes analysis over agent decisions is the evidence an examiner will ask for. Both agencies have signaled the gap is temporary — the OCC has said the agencies plan to issue a request for information addressing model risk management and specifically banks' use of generative and agentic AI. Institutions building this now are building the artifact the next regime will ask for, and it complements rather than replaces existing AML model validation work.

What QA Over Automated Decisions Has to Add

Sampling an agent's dispositions is not the same exercise as sampling an analyst's, and four things change.

The auto-close population needs its own sampling plan, sized independently. The Community Federal order is the cautionary case: the failure was not bad human judgment, it was a high-volume automated disposition path nobody sampled.

Data and pipeline quality control has to be in scope. Automated dispositioning amplifies upstream data defects silently and at scale, which is the UBS lesson applied forward.

Consistency testing becomes possible in a way it never was for humans. Agents can be evaluated on identical or near-identical cases to measure decision stability, which is impractical to test in people and is genuinely stronger evidence.

And override rates work as a two-way signal. Persistently high analyst overrides of agent recommendations indicate miscalibration. Persistently near-zero overrides may indicate rubber-stamping, which is its own finding.

One artifact is worth building even though it is expensive: a held-out re-investigation set, where alerts are re-worked blind by reviewers who cannot see the original disposition. It is the only thing that supports a claim about false negatives, and it must be excluded from any tuning to stay valid.

Where QA Cannot Help

QA reviews alerts that fired. It says nothing about activity that never generated an alert. That is below-the-line testing and threshold tuning — a different control with a different method. Any program that implies QA covers detection failure is describing something it does not do.

Inter-rater reliability is the variable almost nobody measures. If two QA reviewers score the same case differently, the error rate is measuring the reviewers rather than the analysts. Periodic calibration sessions where all reviewers score an identical set, with divergence measured and the rubric revised where they disagree, usually reveal that rubric ambiguity — not analyst error — was the real finding.

And QA cannot cure a resourcing problem. UBS had four employees, none with prior AML experience, reviewing more than 100 unique reports. Canaccord had a daily report its own staff called too long to read. In both cases better QA would have documented the failure sooner without preventing it. That is the honest boundary of the control, and it is also where changing the capacity equation matters more than tightening the measurement of it.

Where Sphinx Fits

Sphinx's agents produce a documented reasoning trail for every disposition through the Interpretable Agentic Framework, which is what makes QA over automated decisions tractable rather than theoretical. A reviewer can see which data the agent retrieved, what it concluded from it, and where the conclusion came from — the same things a QA scorecard asks of a human analyst.

That does not exempt anything from sampling. If agents are closing alerts, the closed population needs its own plan, its own scorecard, and its own held-out set. The institution owns the outcome regardless of what produced it, and the reason a source-linked reasoning trail matters more than the automation question itself is that it is the only thing that makes the sampling meaningful.

Frequently Asked Questions

Does an AML QA program satisfy the independent testing requirement?

No. QA is a second-line control owned by compliance, reviewing whether the first line applied policy correctly. Independent testing is the fourth pillar — it assesses the adequacy of the program as a whole, including QA itself, and reports to the board or a board committee made up primarily of outside directors. The FFIEC also bars persons involved in conflicting BSA functions from performing independent testing, which means a QA team that writes procedures cannot serve as the independent tester.

What is the required AML QA sample size?

There isn't one. No U.S. regulator publishes a benchmark AML QA error rate or a minimum sample size, and the FFIEC makes sampling risk-based and discretionary. Any "industry standard 5 percent" figure is convention rather than regulation. Size samples to the confidence you need per stratum — precision in attribute sampling depends on absolute sample size, not on the percentage of the population reviewed.

Should QA sample escalated alerts or closed alerts?

Closed alerts. Escalations already receive second-level scrutiny and often a SAR decision. False negatives — the alerts that were wrongly closed — exist only in the closed population, and false negatives are what examinations are actually about.

How does QA change when AI agents make dispositions?

Four things. The auto-close population needs its own sampling plan sized independently, because that is where the OCC found a bank auto-closing a very high percentage of ingested alerts. Data and report-generation quality control has to be in scope, because automation amplifies upstream defects silently. Consistency testing on near-identical cases becomes possible and is stronger evidence than anything available for human reviewers. And analyst override rates should be tracked in both directions — high overrides suggest miscalibration, near-zero overrides may suggest rubber-stamping.

Does SR 11-7 still govern AML models?

No. SR 26-2, issued April 17, 2026, superseded SR 11-7 and rescinded SR 21-8, the BSA/AML-specific model risk statement. SR 26-2 also places generative and agentic AI outside its scope entirely and excludes deterministic rule-based processes from the model definition. The underlying legal obligations under 12 CFR 21.21 and 31 CFR 1020.320 are unchanged, which makes an institution's own QA the primary defensible control over automated dispositions until the agencies issue new guidance.

Get Your Free AI Compliance Handbook

What compliance leaders need to know about AI-driven fraud, autonomous laundering, and how your team can
fight back.
Submit
Thank you! Your submission has been received!
Something went wrong while submitting the form. Please try again.