Automated Traceability: Moving from Search Results to Binary Decisions

Automating traceability link recovery through classification

2017-08-02
Chris Mills
Summary
Problem
Method
Results
Takeaways
Abstract

This paper proposes a novel approach to Automating Traceability Link Recovery (TLR) by reframing it as a binary classification problem rather than a traditional Information Retrieval (IR) task. By leveraging machine learning models like Random Forest, the method achieves a high recall (up to 92.7%) in identifying valid links between software artifacts.

TL;DR

Software Traceability Link Recovery (TLR) is traditionally treated as an Information Retrieval (IR) problem, leaving developers to sift through "top-K" results. This paper proposes a paradigm shift: treating TLR as a binary classification task. By combining similarity scores with document statistics and Query Quality (QQ) metrics, the proposed model can automatically label links as valid or invalid, achieving up to 92.7% recall and significantly reducing manual verification effort.

The Bottleneck of Manual Verification

In the software lifecycle, linking requirements to source code or test cases (TLR) is vital for impact analysis. However, current SOTA IR methods (like VSM or LSI) provide ranked lists. The problem? The stakeholder is still the bottleneck. A human must inspect every suggested link to prune false positives. When datasets scale, the manual effort required to find the "needle in the haystack" makes traceability unsustainable for large-scale agile projects.

Methodology: The Classifier's Lens

The core insight of this work is that a link's validity isn't just about textual similarity; it's about the context of the artifacts. The author extracts three types of features to feed into machine learning classifiers:

  1. IR Ranking: Scores from VSM, BM25, and language smoothing models (Jelinek Mercer/Dirichlet).
  2. Query Quality (QQ): Features that determine if an artifact is a "bad query." If two documents are dissimilar, is it because the link is invalid, or because the requirement was poorly written?
  3. Document Statistics: Basic metrics like vocabulary size and term overlap percentages.

Overall Workflow

Tackling Class Imbalance

Valid links are rare (averaging only ~8% of all possible pairs). To prevent the model from simply predicting "invalid" for everything, the author tests two strategies:

  • SMOTE: Generating synthetic examples of valid links.
  • Undersampling: Reducing the number of invalid link examples in the training set.

Critical Results & Performance

The author evaluated several algorithms, including J48, Naive Bayes, and Random Forest.

Performance Comparison Table

Key Insights from the Benchmarks:

  • Random Forest is the winner: It consistently showed the best balance between finding valid links (TPR) and avoiding false alarms (FPR).
  • The Precision-Recall Trade-off: Using SMOTE led to an incredibly low FPR (0.017), meaning the links it did find were almost certainly correct. Undersampling caught more links (92.7% recall) but introduced more noise (12.2% FPR).

Critical Analysis & Conclusion

This work is a significant step toward "Zero-Touch Traceability." By reformulating the problem, it moves the software engineering community closer to tools that don't just suggest links but actually maintain them.

Limitations:

  • Dependency on Labeled Data: Classification requires ground truth. For a brand-new project with no history, a "Cold Start" problem exists.
  • Feature Engineering: The current features are relatively "shallow" (word counts, basic IR). Incorporating semantic embeddings (like CodeBERT) could likely drive the FPR even lower.

Takeaway: The future of software engineering tools lies in automation that removes the human from the loop of trivial verification. This paper provides the foundational framework for that transition in the realm of traceability.

Find Similar Papers

Try Our Examples

  • Find recent papers from 2020-2024 that utilize Deep Learning or Transformers (e.g., BERT) to solve the Traceability Link Recovery classification problem.
  • Which study first introduced the use of Query Quality metrics in software engineering IR tasks, and how has it been integrated into automated traceability since then?
  • Explore how binary classification for traceability has been applied to cross-project scenarios where no labeled training data is available for the target system.
Contents
Automated Traceability: Moving from Search Results to Binary Decisions
1. TL;DR
2. The Bottleneck of Manual Verification
3. Methodology: The Classifier's Lens
3.1. Tackling Class Imbalance
4. Critical Results & Performance
5. Critical Analysis & Conclusion