Tuning Screening and Monitoring: A False-Positive Reduction Methodology
Every firm wants fewer false positives, and the fastest way to get them is to tighten matching until the alerts stop — which also stops the true ones, invisibly. This paper sets out a methodology for reducing alert volume while demonstrating that detection has not degraded.
Regulators do not expect zero false positives. They expect documented tuning — evidence that the configuration was chosen deliberately, tested, and is reviewed. That framing matters, because it means the objective is not the lowest possible alert count.
The objective is the alert population the team can investigate properly, with demonstrated detection of the things that matter.
Why Alert Volume Is a Control Problem
The chain is short. High volume exceeds capacity. Exceeding capacity produces shortcuts — pattern-matched closures, thin rationales, and backlogs. Backlogs delay suspicious reports. Late reports are a regulatory failure in themselves, and a backlog is one of the first things an inspector looks for.
So a tuning programme is not a cost-reduction exercise dressed as compliance. It is a control improvement that happens to reduce cost.
Segment Before You Tune
The most common error is tuning globally. A single threshold across a mixed book is always wrong somewhere: too loose for the low-risk majority, too tight for the segment that actually carries risk.
Segment first — by customer type, risk rating, product, jurisdiction and channel — then tune within segments. A retail customer in a domestic corridor and a corporate customer with layered offshore ownership should not share a rule set. Segmentation also makes the change defensible: "we relaxed matching for this low-risk cohort, on this evidence" is a supervisable statement in a way that "we raised the threshold" is not.
The Screening Side: Match Logic
Screening false positives come from a small number of recurring causes:
- Common names, especially where a name is frequent in a large population.
- Transliteration variance between scripts, where a single underlying name has many valid Latin renderings.
- Insufficient secondary identifiers. A name-only match is weak; name plus date of birth plus nationality is much stronger. Much of the false-positive problem is a data completeness problem.
- List quality — duplicated or poorly structured entries generating multiple hits for one party.
Addressing these is usually more productive than moving the threshold. Capturing date of birth reliably at onboarding will reduce false positives more, and more safely, than a blanket tightening.
The Monitoring Side: Scenario Design
Transaction monitoring false positives tend to come from scenarios that encode an assumption which does not hold for part of the book — a velocity rule calibrated on retail behaviour applied to a corporate treasury flow, or a threshold set in one currency applied across several.
Diagnose before tuning. For each high-volume scenario, sample closed alerts and classify why they were not suspicious. If a single reason accounts for most of them, the scenario has a design fault, not a threshold problem — and changing the threshold will suppress genuine hits along with the noise.
False-Negative Testing: the Part Firms Skip
This is the difference between tuning and degradation. Every reduction must be paired with evidence that detection held.
- Known-entity testing. Maintain a test set of names — sanctioned parties, PEPs, RCAs, including transliterated and hyphenated variants — and run it through the live configuration after every change.
- Above-and-below testing. For monitoring thresholds, sample activity just below the new threshold and review it manually. If suspicious activity is sitting there, the threshold is wrong.
- Historical replay. Run previously-filed reports back through the tuned configuration. Anything that no longer alerts is a finding.
Firms measure false positives because they are visible and expensive. False negatives are invisible — which is exactly why they need deliberate testing rather than monitoring.
Governance and Evidence
Treat every tuning change as a controlled change with: a documented rationale, pre- and post- volume analysis, the false-negative test result, named approval at an appropriate level, and a scheduled review. Supervisors rarely object to a well-evidenced threshold. They object to one nobody can explain.
A tuning log that records what changed, why, who approved it and what testing was done answers most of the questions an inspection will raise on this subject.
Machine Learning: What It Solves and What It Does Not
Machine learning is routinely offered as the answer to false positives. It can genuinely help, in specific places, and it introduces obligations that firms frequently underestimate.
Where it helps. Alert triage and prioritisation — ranking a queue so analysts reach the likeliest cases first. Entity resolution — deciding whether two similar records are the same party, which is a pattern problem classical matching handles poorly. Anomaly detection as a supplement to rules, surfacing behaviour no scenario anticipated.
Where it creates exposure. Auto-closing alerts without human review is the obvious one, and most supervisors will want to understand the control around it very carefully. Beyond that, two structural issues: a model trained on historical dispositions learns your existing biases, including anything previously closed incorrectly; and a model whose decisions cannot be explained in plain terms is difficult to defend when an inspector asks why a specific alert was deprioritised.
The practical test is whether you could explain a model-influenced decision to a supervisor without reference to the model’s internals. If not, use it to order work rather than to decide it.
Running Tuning as a Function, Not a Project
Most tuning happens as a one-off remediation after alert volumes become intolerable. That produces a step change, a period of stability, and then the same problem two years later as the book changes underneath the configuration.
Treating it as an ongoing function requires four things:
- Standing metrics. Alert volume, disposition split and investigation time per scenario and per segment, reviewed on a regular cycle rather than when something goes wrong.
- A defined change process. Proposal, impact analysis, false-negative test, approval, implementation, post-change review — the same discipline applied to any control change.
- Clear ownership. Someone owns the configuration. Where tuning is split between compliance, operations and a vendor, nobody holds the overall picture and changes accumulate without a coherent rationale.
- An annual independent review. Internal audit or a third party testing whether the configuration still matches the risk assessment. This is increasingly an explicit supervisory expectation rather than good practice.
The output of a mature function is not a lower alert count. It is the ability to answer, for any scenario: why is it configured this way, when was that decided, what testing supports it, and when will it next be reviewed.
What Good Looks Like
A mature programme can state, for each segment: the current configuration, when it was last reviewed, the alert volume and disposition split, the false-negative test result at the current setting, and the rationale for the last change.
That is a materially different posture from "our false-positive rate is down 40%" — which, on its own, is as consistent with degraded detection as with improved tuning.
Tune With Evidence, Not Guesswork
Segmented configuration, per-scenario diagnostics, and false-negative testing built into the change process — with the tuning log an inspector will ask to see.
