Independent editorial research

Rules, Machine Learning, and Explainability in AML Systems

A neutral comparison of rules and machine-learning methods, the explanations each can support, and the evidence buyers need before approving an AML use case.

By AML Tech Reviews Editorial TeamPublished Updated

A rule can say, “alert when a defined threshold is crossed.” A model may say, “this activity differs from the learned pattern.” Both statements leave important questions unanswered: which risk is being addressed, which data was used, why this customer was selected, and what evidence supports the decision?

Rules and machine learning are methods, not control objectives. Either can be useful or weak depending on data, design, intended use, testing, implementation, and oversight. A hybrid system can combine explicit rules, statistical features, peer comparisons, supervised models, anomaly detection, and investigator judgment.

What rules do

A rule applies defined logic to inputs. It can use a single threshold, several conditions, aggregation over time, segmentation, or a sequence of events. Rules are often easy to trace because the conditions are explicit. They are not automatically simple. Thousands of interacting rules, exceptions, thresholds, and segments can become hard to understand and maintain.

Rule coverage depends on the risk hypothesis and data. A clear rule can consistently miss activity that falls outside its assumptions. A rule can also generate unnecessary work when a threshold is copied across populations with different normal behavior.

The FCA’s Financial Crime Guide asks firms to understand monitoring rules, thresholds, capabilities, and limitations, and to update rules for new trends. That supports documented rationale and review. It does not require a rules-only architecture.

What machine learning does

Machine learning estimates patterns or relationships from data. In supervised learning, examples with labels are used to predict an outcome. In unsupervised or anomaly methods, the system can identify unusual observations or groups without the same labelled target. Models can rank, classify, cluster, or create features used by another control.

The intended use matters. A model that prioritizes an existing alert queue has a different risk from one that replaces a detection scenario or automatically closes cases. Training labels can reflect past rules and investigator decisions, including their gaps. An anomaly is not the same as suspicious activity, and a low model output is not proof that activity is safe.

The Wolfsberg AI and machine-learning principles identify legitimate purpose, proportionate use, design and technical expertise, accountability and oversight, openness, and transparency. They also call for attention to fair, effective, explainable outcomes and stable performance. These are practitioner principles, not a substitute for applicable law or an institution’s model-risk framework.

Explainability has several audiences

An investigator may need to know which transactions, customer attributes, counterparties, or behaviors drove an alert. A validator may need model design, features, data lineage, segment performance, limitations, and sensitivity. A compliance owner may need coverage against assessed risks. An auditor may need to reproduce the historical output. A customer-facing process may carry separate explanation duties under applicable law.

One graphic cannot meet all of those needs. Feature contribution methods can describe how inputs influenced an output, but they may be approximations and do not prove that a model is correct. A rule trace can show fired conditions while leaving the threshold rationale unsupported. Explanation should connect the output to source data, method, version, intended use, and decision path.

The Wolfsberg Part II monitoring statement discusses explainability alongside transition, validation, and the balance between model risk and financial-crime risk. It argues for grounding model design in a comprehensive risk assessment. That makes explainability part of control design, not a report added after deployment.

Evidence to test

For rules, request the risk rationale, population, data fields, logic, parameters, exceptions, dependencies, historical changes, test results, and retirement criteria. Ask how overlapping rules are handled and how a threshold change affects detection and workload.

For machine learning, request the intended use, decision role, training and validation populations, time periods, label definitions, feature lineage, exclusions, data-quality controls, algorithm choice, parameter selection, segment results, uncertainty, limitations, monitoring, override, retraining, and rollback. If the supplier cannot disclose proprietary implementation detail, the buyer still needs enough evidence to govern the use and test its behavior.

Run both methods on representative, labelled cases and ordinary activity. Include data shifts, missing fields, new customer segments, outliers, duplicates, and boundary cases. Measure detection and false alerts on the test set, stability, segment differences, investigator interpretation, and the ability to reproduce outputs. Compare the proposed method with the current control or a meaningful baseline.

The Bank of England and FCA’s 2022 machine-learning survey reported respondents’ concerns about bias, data quality and structure, explainability, inaccurate predictions, governance, and outsourcing. A survey describes respondent experience; it does not establish the performance of a particular AML model.

Procurement implications

Buy the governance capability with the analytic capability. Contracts and architecture should support access to data definitions, version history, validation evidence, output reconstruction, performance monitoring, material-change notice, incident support, and usable export. Clarify which party owns features, labels, tuning, validation support, and intellectual property created from customer data.

Score the use case, risk coverage, data readiness, evidence, explanation, validation, human oversight, change control, resilience, integration, and total operating cost separately. A supplier’s generic model result is not a proof-of-concept result. Test the configured use on the institution’s representative data and workflow.

Plan transition. Parallel operation can reveal differences, but the comparison needs an agreed method for examining cases found by only one system. A new method may identify useful activity the old process missed; it may also create a new blind spot. Define approval and rollback conditions before production.

Evidence limitations

True financial-crime labels are incomplete. Reported cases reflect what prior controls and people found, while unreported activity remains unknown. Synthetic examples can test behavior but not prevalence. Model performance can change as customers, products, data, and threats change. Rules can also decay when their assumptions no longer hold.

Explainability does not guarantee fairness, accuracy, or effectiveness. Complexity does not prove value, and simplicity does not prove control. The defensible choice is the method whose purpose, data, behavior, limitations, and oversight the institution can demonstrate for the specific use.