AML Internal Audit: A Third-Line Testing Checklist
There is a difference between checking that work was done and testing whether a control works. First-line checklists do the first. Third-line audit exists to do the second, and the distinction is what makes it independent rather than duplicative.
Scoping by risk, not by calendar
Audit universes that allocate coverage evenly across the programme waste effort. Scope should follow the enterprise-wide risk assessment, with three modifiers: areas where inherent risk is highest, areas where controls are newest or recently changed, and areas where prior findings remain open.
A three-year rolling cycle covering every component, weighted so high-risk components are tested annually, is the common and defensible pattern. What is not defensible is a cycle that has never tested transaction monitoring rule logic because it is technical, which is a surprisingly frequent gap.
What to test, and how
Governance and independence
Test whether the MLRO has genuine authority: reporting line, budget control, and whether they can escalate without going through a revenue owner. Read board and committee minutes for evidence of challenge, not just presentation. Management information that reports volumes without reporting risk is a finding in itself.
Risk assessment
Do not re-perform it — test it. Is it current? Does it cover products launched since the last version? Are control effectiveness ratings supported by test results, or asserted? Trace three material risks to the controls that address them.
Customer due diligence
Sample from the full population, stratified by risk tier. For each file test completeness against policy, whether the risk rating applied matches the methodology, whether EDD triggers fired where they should have, and whether EDD that was triggered was actually completed. That last test finds more issues than the other three combined.
Screening
Test configuration, not just output. Which lists are loaded, at what update frequency, with what matching threshold, and against which fields? Then test it empirically: inject known test entities — including name variants, transliterations and partial matches — and confirm they alert. A screening system nobody has probed is a system nobody knows the sensitivity of. See sanctions list screening.
Transaction monitoring
This is where audit most often retreats into reading documentation. Resist that. Three tests are feasible without deep technical skill. First, coverage: list the typologies in the risk assessment and identify which have no corresponding rule. Second, parameter provenance: for a sample of rules, ask why the threshold is what it is and request the tuning analysis. Third, above-the-line and below-the-line testing: review a sample of alerts generated, and — more informative — a sample of activity that sat just below threshold, to see whether anything obviously suspicious escaped. See rule tuning.
Alert and case handling
Test disposition quality, not just timeliness. Can an independent reader reconstruct why the alert was closed from the note alone? Test the escalation path: were cases meeting SAR criteria escalated, and within what period? Check ageing and any backlog.
Suspicious activity reporting
Sample filed reports for quality and timeliness, and — harder and more valuable — sample non-filed escalations that were closed without a report, to test whether the decision not to file was reasoned. See SAR filing.
Training
Completion rates are the weakest possible measure. Test whether content is role-appropriate, whether assessment results are retained, and whether the people handling the highest-risk work received training matched to it.
Sampling that withstands challenge
Three rules. Select randomly from the full population — never accept a list prepared by the area under audit. Stratify by risk so high-risk files are over-represented relative to their share, and state the stratification. And size the sample so a conclusion is supportable; if you cannot justify the size statistically, describe it as indicative rather than conclusive and say so in the report.
Where a sample produces exceptions, expand it before concluding. A finding based on two exceptions in twenty-five is materially weaker than one based on an expanded sample that establishes a rate.
Grading and reporting
Grade issues on the risk they create, not on how difficult they are to fix. A high-rated finding should mean a control does not work or does not exist for a material risk — not that a procedure document is out of date. Grade inflation destroys the audit committee's ability to prioritise, and grade deflation is worse.
Every finding needs a named owner, an agreed action and a date. Track them to closure and validate closure by re-testing, not by accepting an assertion that the work was done. Report ageing of overdue findings to the audit committee — a pattern of slipped dates is itself a governance finding.
Where audit adds the most value
In practice, the highest-value tests are the ones the first line cannot perform on itself: monitoring coverage against typologies, below-the-line testing, and non-filed escalation review. These are also the tests examiners increasingly expect to see. A third-line function that tests policy adherence but has never opened the monitoring configuration is not providing the assurance the board believes it is receiving.
Related: preparing for an AML examination covers what happens when an external party runs equivalent tests.
Independence in Practice
Independence is easy to claim on an organisation chart and harder to demonstrate in operation. Three tests matter.
Who sets the scope
If the area under audit materially shapes what gets tested, independence is compromised regardless of reporting lines. Scope should follow the risk assessment and prior findings, with management consulted rather than deciding.
Who selects the sample
A sample supplied by the audited function is not a sample. Audit should draw from the full population itself, and should be able to demonstrate how.
Who grades the finding
Negotiating a finding down from high to medium because remediation would be expensive is the point at which the third line stops functioning. Management may disagree, and that disagreement should be recorded alongside the finding rather than resolved by changing it.
For smaller firms without a standing audit function, the UK's Regulation 21(1)(c) obligation can be met by an appropriately skilled person who is independent of the programme — an external reviewer, or someone from elsewhere in the group. What does not satisfy it is the compliance function reviewing itself and calling the result an audit.
Using Data in AML Audit
Sampling tells you about the files you looked at. Population-level analysis tells you about the book, and it is increasingly what examiners expect to see in audit workpapers.
Four analyses are feasible without specialist tooling. Below-the-line testing: extract activity sitting just under each monitoring threshold and review a sample — this is the most direct evidence of whether thresholds are calibrated or merely set. Coverage mapping: list the typologies in the risk assessment against the rules in production and identify which typologies no rule addresses. Disposition timing: plot time-to-close by analyst; unusually fast closures cluster around capacity pressure. CDD completeness: query the whole customer base for missing mandatory fields by risk tier rather than inferring completeness from a sample of thirty.
Each of these turns a qualitative opinion into a measured finding, and each is repeatable year on year — which makes the trend as informative as the result.
Reporting That the Committee Can Act On
An audit report that lists findings without context leaves the committee unable to prioritise. Three additions change that.
State the coverage, not just the result
Say what was tested, what was not, and why. A clean report on a narrow scope reads very differently from a clean report on a broad one, and the committee cannot distinguish them unless told.
Separate design from operation
A control that is poorly designed and a control that is well designed but inconsistently performed require different remediation and different owners. Reports that conflate the two produce action plans that fix the wrong thing.
Show the trend
Repeat findings are the most important signal a third line produces. A finding appearing for the second or third cycle is no longer a control weakness; it is a governance one, and it should be presented that way rather than renumbered and restated.
Finally, report the ageing of overdue actions every cycle, with the original and revised dates both visible. A pattern of quietly extended deadlines is difficult to see in a single report and obvious across four.
Controls You Can Test, Not Just Read About
One Constellation exposes the configuration, tuning history and disposition quality that third-line testing depends on — so audit can test the control itself.
