Agilkia: Bridging the Gap Between Real-World Usage and Automated Testing

Identifying and Generating Missing Tests using Machine Learning on Execution Traces

2020-08-01
Mark Utting, Bruno Legeard, Frédéric Dadeau, Frédéric Tamagnan, Fabrice Bouquet
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces Agilkia, a toolkit that uses Machine Learning to identify testing gaps by comparing customer execution traces with test traces via clustering. It further leverages predictive models like Random Forests to automatically generate new, realistic test cases that cover previously untested user behavior patterns.

TL;DR

The "Agilkia" framework utilizes Machine Learning to mine system logs and identify where your tests are failing to match reality. By clustering user traces and comparing them to existing test suites via 2D visualization, it identifies "untested" behaviors and uses predictive models (like Random Forests) to automatically synthesize new, realistic regression tests.

Background: The QA Bottleneck

Despite the rise of DevOps, testing remains a friction point. Most test suites are "stereotyped"—they follow the developer's happy path but ignore the messy, idiosyncratic ways actual customers use an API. This paper addresses the "What are we missing?" question by treating execution traces as the ground truth of system usage.

Problem & Motivation

The authors identify a critical gap: Operational data (logs) is rarely used for validation.

  1. Complexity: Symbolic model-based testing is mathematically rigorous but often too hard for industry engineers to maintain.
  2. Lack of Realism: Hand-written tests rarely cover the combinatorial explosion of actions users perform. The insight here is simple yet powerful: If we can cluster similar user behaviors, we can see exactly which clusters have zero test coverage.

Methodology: From Logs to Predictive Models

The Agilkia pipeline consists of four distinct stages:

1. Trace Preprocessing & Encoding

Raw logs (CSV/Json) are split into individual user sessions and encoded. The paper uses a Bag-of-Words approach, treating each API call as a "word" and the session as a "document." This creates a numerical vector that reflects the frequency of actions.

2. Clustering for Behavior Mapping

Using the MeanShift algorithm, Agilkia automatically groups similar sessions. Unlike K-Means, MeanShift doesn't require pre-defining the number of clusters, making it ideal for discovering unknown usage patterns.

3. Visual Gap Analysis

The authors use Principal Component Analysis (PCA) to project high-dimensional traces onto a 2D map. PCA Visualization of Traces Fig: Visualization of customer traces. By overlaying existing tests on this map, engineers can visually identify "white spaces" where no tests exist.

4. Automated Test Generation

Once a missing cluster is identified (e.g., "users who scan 20+ items but then cancel"), Agilkia trains a supervised ML model (Random Forest or Decision Tree) on those specific traces. The model learns to predict the "Next Action" given a prefix, allowing the SmartSequenceGenerator to roll out new, valid test flows.

Experiments & Results

The framework was validated on three distinct systems:

  • Scanner Simulator: Identified that manual tests missed 6 out of 13 behavior clusters.
  • Bus Tracking System: Generated tests that covered 33.75% of behavior from just a few traces, successfully identifying error-handling paths.
  • Supply Chain Web Service: Achieved 60% behavioral coverage with a small set of systematically generated tests.

F1 Score Comparison Table: Random Forest and Decision Trees outperformed other classifiers in predicting the next test action, achieving F1 scores > 0.95.

Critical Insight & Conclusion

The true value of Agilkia isn't just in generating tests—it's in the quantification of coverage. Instead of saying "we have 80% code coverage," engineers can now say "we cover 80% of actual customer behavior patterns."

Limitations

  • Data Dependencies: The current generation focuses on action sequences but struggles with complex input data values (e.g., specific strings or IDs).
  • Trace Quality: The method relies on the availability of high-quality session-tagged logs.

In summary, Agilkia shifts the testing paradigm from "what could happen" to "what is happening," providing a practical, ML-driven path to more resilient software.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use Deep Learning sequence models like LSTMs or Transformers specifically for web API test sequence generation from logs.
  • Which paper originally proposed the MeanShift algorithm for clustering, and what are its specific advantages over K-Means in the context of high-dimensional software trace analysis?
  • Explore how the Agilkia framework's clustering-based test identification could be applied to mobile UI testing or microservices mesh telemetry.
Contents
Agilkia: Bridging the Gap Between Real-World Usage and Automated Testing
1. TL;DR
2. Background: The QA Bottleneck
3. Problem & Motivation
4. Methodology: From Logs to Predictive Models
4.1. 1. Trace Preprocessing & Encoding
4.2. 2. Clustering for Behavior Mapping
4.3. 3. Visual Gap Analysis
4.4. 4. Automated Test Generation
5. Experiments & Results
6. Critical Insight & Conclusion
6.1. Limitations