Two-Layered AML: Balancing Extreme Scalability with Imbalance-Resistance

Scalable and Imbalance-Resistant Machine Learning Models for Anti-money Laundering: A Two-Layered Approach

2020-01-01
Pavlo Tertychnyi, Ivan Slobozhan, Madis Ollikainen, Marlon Dumas
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a two-layered machine learning framework for Anti-Money Laundering (AML), utilizing Logistic Regression and Gradient Boosting (CatBoost) to detect illicit financial behavior. The approach achieves state-of-the-art efficiency by combining rapid filtering with high-complexity classification on a real-world dataset of 330,000 customers.

TL;DR

Financial institutions process millions of transactions per minute, making traditional manual or rule-based monitoring obsolete. This paper presents a two-layered machine learning architecture that first "filters the noise" using simple Logistic Regression and then "hunts the signal" using CatBoost with complex sequence-based features. The result? A system that is 40% faster to train and more accurate than standard single-classifier SOTA models.

The Tension: Scalability vs. Accuracy

In the world of Anti-Money Laundering (AML), we face two brutal realities:

  1. Extreme Imbalance: Only about 0.004% of customers are actually illicit.
  2. Computational Bottleneck: Extracting "deep" features (like transaction sequences or network statistics) for 100% of a bank's customers is a recipe for infrastructure collapse.

Previous works often relied on random undersampling or oversampling (SMOTE). However, the authors argue that random sampling ignores a critical insight: most customers are obviously innocent. By ignoring this, we waste cycles computing complex features for customers who would never be flagged anyway.

Methodology: The Two-Layered Filtration

The core innovation is a hierarchical pipeline designed to optimize the "computational budget."

Layer 1: The Massive Scale Filter

  • model: Logistic Regression (LR).
  • Features: Simple Descriptive Statistics (mean/std of amounts) and demographics.
  • Goal: Extreme Recall. The authors tuned the threshold to ensure that 99% of illicit cases pass through, while filtering out ~48% of the innocent majority.

Layer 2: The Precision Hunter

  • Model: CatBoost (Extreme Gradient Boosting).
  • Features: Complex engineering including Sequence-based features (Generative log-odds using Markov Chains) and Counterparty statistics.
  • Execution: This layer only processes the 52% of customers flagged by Layer 1.

Final Classification Model Architecture

Feature Engineering: Markov Chains for Transactions

One of the most impressive technical contributions is the use of Generative Log-Odds features. Instead of just looking at transaction amounts, the system models the transition between states (e.g., Inbound -> Outbound -> Cash Withdrawal). By calculating the log-probability of a sequence being illicit vs. non-illicit using the Markov property, they compress time-series dynamics into a single, high-signal feature for the Gradient Boosting model.

Experimental Results

The framework was tested on a massive dataset of 330,000 customers and 51 million transactions from three different countries.

Performance Boost

As shown in the table below, the Two-layered model consistently outperformed single-layer CatBoost and Random Forest across different risk thresholds.

Overall Results Comparison

Efficiency Gains

The hardware-efficient nature of Layer 1 means the system doesn't have to extract the 400+ complex features for nearly half the population. This resulted in a training time reduction from 148 minutes to 89 minutes.

Execution Time Table

Critical Insight & Future Directions

The paper proves that in production environments, the architecture of the pipeline is as important as the model itself. By treating the first layer as a "knowledge-driven undersampler," the authors effectively solve both scalability and imbalance in one move.

Limitations: The model is purely supervised. It can only find what experts have found before. To catch new money laundering schemes, future iterations must integrate Unsupervised Anomaly Detection to flag patterns that don't fit the "norm" but haven't been labeled yet.

Conclusion: For practitioners building AML systems, this paper is a blueprint for moving away from heavy "single-shot" models toward lean, tiered architectures that respect the constraints of real-world banking infrastructure.

Find Similar Papers

Try Our Examples

  • Find recent papers that combine anomaly detection (unsupervised) with supervised learning for anti-money laundering to address the "previously identified patterns" limitation.
  • Which 2024-2025 SOTA methods utilize Graph Neural Networks (GNNs) for AML ego-network analysis, and how do they compare in scalability to this two-layered approach?
  • Explore how generative AI and Large Language Models (LLMs) are being used to process the "additional documents" (invoices, contracts) mentioned in this paper's future work section.
Contents
Two-Layered AML: Balancing Extreme Scalability with Imbalance-Resistance
1. TL;DR
2. The Tension: Scalability vs. Accuracy
3. Methodology: The Two-Layered Filtration
3.1. Layer 1: The Massive Scale Filter
3.2. Layer 2: The Precision Hunter
4. Feature Engineering: Markov Chains for Transactions
5. Experimental Results
5.1. Performance Boost
5.2. Efficiency Gains
6. Critical Insight & Future Directions