One Constellation
White Paper

Selecting an AML/KYC Platform: A Buyer’s Evaluation Framework

Most AML platform selections are decided on a feature matrix, and most feature matrices are answered "yes" by every vendor in the process. This paper sets out an evaluation framework built around demonstration rather than declaration — what to ask, what to watch, and which answers actually differentiate.

Published: September 2026 Section: White Papers Read time: ~12 minutes
Executive Summary
A feature checklist cannot separate AML platforms, because every serious vendor answers yes to every line. What differentiates them is how a capability works: whether risk scoring is configurable by you or by the vendor, whether screening reaches beneficial owners and connected parties or only the named customer, whether an alert’s reasoning can be reconstructed months later, and whether the data model survives the structures you actually onboard. This paper proposes six evaluation dimensions, a weighted scoring method, and a set of live demonstrations to require — using your data, not the vendor’s.

Compliance technology selections fail in a predictable way. A requirements matrix is circulated, every vendor returns it fully ticked, the shortlist is decided on price and presentation, and the gaps surface in month four of implementation — when the risk model cannot express the firm’s actual policy, or the screening scope turns out to exclude the parties that mattered.

The failure is not diligence. It is that the questions asked were answerable in the abstract. This paper reframes the evaluation around things a vendor must show rather than assert.

Why Feature Matrices Do Not Discriminate

Consider a typical requirement: "The platform must support PEP screening." Every vendor says yes. The answer conceals the questions that decide whether the control works:

  • Does it screen the customer only, or also beneficial owners, appointed persons and connected parties?
  • Does it cover domestic and international-organisation PEPs, or foreign only?
  • Does it include relatives and close associates, and if so from what source?
  • How is match confidence tuned, who can tune it, and is the change audited?
  • What happens when a customer becomes a PEP after onboarding?

Five firms answering "yes" to the requirement may differ completely across these. The framework below is built to surface that difference.

The Six Dimensions

1. Configurability without vendor dependency

The central question: when your risk policy changes, who implements it? If a new jurisdiction risk weighting, a new EDD trigger or a new monitoring scenario requires a vendor change request, your compliance policy is gated by someone else’s release cycle. Ask to see a risk factor added and weighted live, by a non-engineer, during the demonstration.

2. Screening scope and match quality

Scope is the first question and match quality the second. Ask for the false-negative test: supply known PEPs and sanctioned parties, including transliterated and hyphenated names, and confirm they are detected at the configured threshold. Most firms measure false positives because they are expensive and visible; far fewer test what the configuration is failing to catch.

3. The data model behind the customer

Retail banking assumes a customer is a person or a company. Fund structures, trusts and layered corporate ownership break that assumption. Ask the vendor to model one of your more awkward structures — a fund with a management company, a nominee register and an underlying trust — and ask where the risk rating attaches, and whether the same natural person appearing in three structures is one record or three.

4. Evidence and reconstruction

Supervisors test whether a decision can be reconstructed, not whether it was recorded. Take a closed alert in the demo environment and ask to rebuild it: what triggered it, what data the analyst saw, what was considered, why it was closed, who approved it, and what the system state was at that moment. A platform that stores outcomes but not reasoning will fail an inspection regardless of its detection quality.

5. Jurisdictional fit

If you operate across, say, Singapore, Hong Kong and Luxembourg, the platform has to express three supervisors’ expectations without three separate builds. Ask specifically how a single control set carries different thresholds and obligations per entity — for example the S$5,000 occasional transaction trigger under MAS Notice PSN01 alongside EU requirements.

6. Operational load

A platform that detects well but generates more alerts than the team can investigate produces rushed closures and late reports. Ask for realistic alert volumes at your transaction profile, and what tuning is available to you directly.

A Weighted Scoring Method

Score each dimension 1–5 against evidence shown rather than claims made, and weight by what actually carries risk in your firm:

  • Configurability — weight high if your policy changes often or you operate in several jurisdictions.
  • Screening scope and quality — weight high in almost all cases; this is where enforcement concentrates.
  • Data model — weight high for fund administration, trust and corporate services, wealth.
  • Evidence — weight high if you are supervised on-site.
  • Jurisdictional fit — weight by number of regulated entities.
  • Operational load — weight high where team size is fixed and volumes are growing.

Require an evidence note against every score. A dimension scored from a slide rather than a demonstration should be capped at 3.

Demonstrations to Insist On

Five, each on your data:

  1. Onboard a complex structure and show where ownership resolves and where the risk rating attaches.
  2. Add a risk factor and a monitoring scenario live, without vendor engineering.
  3. Run a false-negative test with names you supply, including transliterations.
  4. Reconstruct a closed alert end to end.
  5. Produce the supervisory export — the actual file or report your regulator asks for.

A vendor who cannot do these in a controlled demonstration will not do them under inspection.

Questions That Do Not Differentiate

These consume RFP space and separate nobody. Ask them briefly or not at all:

  • "Do you support sanctions screening?" — everyone does. Ask about scope and tuning instead.
  • "Is the platform cloud-hosted?" — ask about data residency, which is the real question.
  • "Do you have an API?" — ask what is exposed through it and what is not.
  • "How many customers do you have?" — ask for references in your sector and jurisdiction.
  • "Are you ISO 27001 certified?" — useful, but a floor rather than a differentiator.

Commercial Terms That Decide the Next Five Years

Evaluation stops at functionality far too often. Four commercial terms shape the relationship more than any feature:

  • Data ownership and extraction. If you leave in three years, what comes with you and in what format? Customer records, risk ratings, alert history and decision rationale should all be extractable in a structured form. "We will provide a data dump" is not an answer — ask to see the export schema.
  • Price escalation mechanics. Per-check pricing behaves very differently from per-seat as volumes grow. Model your three-year volume plan against the pricing structure, including the enhanced-due-diligence cases that consume multiple checks each.
  • Change request treatment. Establish before signing what counts as configuration (included) and what counts as development (chargeable). Vendors differ enormously here, and it is the single most common source of post-contract friction.
  • Regulatory change commitments. When a supervisor changes a requirement, who implements it and on what timeline? A platform serving your jurisdictions should be tracking those changes; ask for examples of how the last three were handled.

None of these appear on a feature matrix, and all of them will matter more in year three than the capability differences that dominate the selection.

Implementation Risk: What the First Ninety Days Reveal

Selection failures usually surface during implementation rather than in production, and they cluster in three places.

Data migration. Moving existing customers into the new platform exposes every gap in the old one. Records without verified beneficial ownership, risk ratings with no recorded basis, alerts closed without rationale — all of it has to be carried, remediated or explicitly written off. Ask the vendor how migrated records are marked, because a supervisor will want to distinguish a rating the new system calculated from one inherited on trust.

Policy translation. Your written policy has to become configuration. This is where firms discover their policy contains ambiguities that a human analyst resolved by judgement and a system cannot. Budget for policy clarification work, and treat any question the configuration cannot express as a policy defect rather than a platform limitation.

Parallel running. Run old and new together for a defined period and compare outputs. Where they disagree on a risk rating or an alert, understand why before switching off the old system. Firms that skip this step lose the ability to explain a change in alert volume at exactly the moment a supervisor asks about it.

Ask each shortlisted vendor to describe these three phases for a client of your size and profile, with timelines. The quality of that answer is itself a strong signal.

What This Looks Like in Practice

A mid-size fund administrator running this framework will typically find the shortlist reorders between the paper evaluation and the demonstrations. Vendors that scored well on features often score poorly on the data model — because fund structures were an afterthought — and vendors that looked unremarkable on paper sometimes score well on evidence and configurability.

That reordering is the point. The framework is designed to move the decision from what a platform claims to what it demonstrably does.

Put Us Through This Framework

We will run all five demonstrations on your data — a real structure, your risk policy, your names for the false-negative test, and the supervisory export your regulator asks for.

Best AML Software → The True Cost of AML Compliance All White Papers
Scroll to Top