ML-Powered Conformance Checking: Accelerating Process Mining with ACTC
An Alignment Cost-Based Classification of Log Traces Using Machine-Learning
This paper introduces the Alignment Cost Threshold-based Classification (ACTC), a novel approach using Machine Learning (Random Forest and LSTM) to classify log traces based on their proximity to process models. By treating conformance checking as a binary classification task, it achieves high-speed trace evaluation and establishes a theoretical lower bound for process model fitness.
TL;DR
Conformance checking—the process of ensuring real-world event logs match theoretical process models—is notoriously slow. This paper proposes ACTC (Alignment Cost Threshold-based Classification), a method that uses Random Forests and LSTMs to predict if a trace "fits" a model within a certain cost threshold. The result? Near-instantaneous trace evaluation and a proven lower bound for model fitness.
The Scalability Wall in Process Mining
In modern enterprises, event logs are massive. To check if these logs comply with prescribed business processes, researchers use alignments. An alignment maps log activities to model activities, penalizing "skips" or "insertions."
The problem is that finding the optimal alignment is a search problem with high computational complexity. As models grow, calculating fitness for thousands of traces becomes a bottleneck. The authors ask a critical question: Can we replace exact computation with a high-accuracy ML "oracle"?
Methodology: From Sequences to Classes
The authors transform the conformance problem into a binary classification task. By setting a threshold , a trace is considered Positive if its alignment cost is low (close to the model) and Negative if it deviates significantly.
The Two-Pronged Architecture
- Random Forest (RF) + Bag-of-Words (BoW): This approach ignores the order of events but captures the frequency of activities. It serves as a fast, robust baseline.
- Bi-LSTM (Long Short-Term Memory): This deep learning model processes traces as sequences, using embedding layers to capture the temporal relationships and "memory" of process flows.
Figure 1: The experimental pipeline from log traces to ML-based classification.
The Fitness Lower Bound Theorem
One of the paper's strongest contributions is Theorem 1. It mathematically proves that you don't need exact alignment costs for every trace to estimate the overall health of a process. By knowing how many traces fall into the "Positive" class, we can determine a guaranteed lower bound for the model's fitness score.
Experimental Battleground: BPI Challenges
The researchers tested their models on real-world data from the Business Process Intelligence (BPI) Challenges (2012, 2017, and 2019).
Key Performance Metrics
- Accuracy: Both RF and LSTM consistently scored above 90% accuracy across varied datasets.
- Inference Speed: Once trained, the RF classifier could evaluate a trace in 0.03ms, compared to up to 99.89ms for traditional tools like ProM. This is a massive leap for real-time monitoring.
Table 1: Comparing ML predictions (RNN/RF) against exact ProM alignment run-times.
Critical Insight: Training vs. Execution
While the "execution" of the ML model is blazing fast, the authors provide a transparent look at training costs. Training an LSTM can take several hours, whereas a Random Forest takes mere seconds.
The Trade-off:
- LSTM is better for complex, sequence-dependent processes where the order of activities is the defining compliance factor.
- Random Forest is superior for high-throughput environments where the presence/absence of certain activities (BoW) is a sufficient proxy for conformance.
Conclusion and Future Outlook
The paper successfully bridges the gap between formal process logic and statistical machine learning. By proving that a binary classifier can provide a lower bound for fitness, it opens the door for:
- Real-time auditing: Flagging non-compliant traces the moment they occur.
- Regression-based Alignments: Moving from "Is it compliant?" (Classification) to "Exactly how non-compliant is it?" (Regression).
Limitations: The current approach requires a pre-existing "oracle" (exact alignments) to generate the training data. Future work might look into unsupervised or semi-supervised ways to achieve similar bounds without the initial heavy compute.
