Automatic Traceability: The Bridge Between Software Evolution and Impact Analysis

A Literature Review of Automatic Traceability Links Recovery for Software Change Impact Analysis

2020-07-13
Thazin Win Win Aung, Huan Huo, Yulei Sui
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a Systematic Literature Review (SLR) of 33 primary studies focusing on automatic Traceability Links Recovery (TLR) for Software Change Impact Analysis (CIA). It categorizes methods into Information Retrieval (IR), Heuristic, Machine Learning (ML), and Deep Learning (DL) based approaches, identifying a shift toward DL for bridging knowledge gaps between heterogeneous artifacts.

TL;DR

Software systems are in a constant state of flux. Every change to a requirement or a bug fix can trigger a "ripple effect" across code, design, and tests. This paper provides a comprehensive literature review (2012–2019) of Traceability Links Recovery (TLR) techniques. It identifies how the field is moving from manual keyword matching toward intelligent, deep-learning-based systems that can "understand" the relationship between human language and machine code.

The "Knowledge Gap" Problem

The primary friction in Software Change Impact Analysis (CIA) is the Knowledge Gap. Requirements are written in natural language (informal, ambiguous), while source code is written in programming languages (formal, strict syntax).

Traditional Information Retrieval (IR) approaches assume that if two documents share words, they are related. However, the authors argue this is insufficient for complex evolution where "method names" might not appear in "user stories."

Methodology: Mapping the Landscape

The authors categorized the research into four distinct technological generations:

  1. Information Retrieval (IR): Using Vector Space Models (VSM) and Latent Semantic Indexing (LSI) to find textual overlaps.
  2. Heuristics: Adding rules based on change frequency and recency.
  3. Machine Learning (ML): Training classifiers (Random Forest, SVM) to filter the "noise" (false positives) generated by IR engines.
  4. Deep Learning (DL): The current frontier, using Word Embeddings and Recurrent Neural Networks (RNNs) to capture deep semantic meanings rather than just keyword counts.

The Bohner Model Integration

A standout feature of this review is how it maps traceability onto Bohner’s CIA Process Model, which divides impacts into four categories:

  • RIS (Requirements Impact Set)
  • DIS (Design Impact Set)
  • TIS (Test Impact Set)
  • PIS (Program Impact Set)

Bohner CIA Process Model The model demonstrates the iterative nature of identifying impacts across different levels of abstraction.

Key Insight: Direct vs. Transitive Tracing

One of the most valuable findings in this review is the distinction in recovery methods:

  • Direct Tracing: Works when Source A and Target B are similar.
  • Transitive Tracing: Essential when A and B have no common vocabulary. If Requirement A relates to Bug Report C, and Bug Report C relates to Code B, a link is inferred.

Traceability Recovery Methods

Experimental Results & Trends

The review highlights a significant experimental bias:

  • 82% of studies use open-source projects.
  • Only 9% involve real-world industrial case studies.
  • SOTA Achievement: Deep learning approaches are proving superior at handling "polysemy" (words with multiple meanings) and "synonymy," which are the death knells of traditional IR.

Research Methods and Datasets

Critical Analysis & Future Outlook

While the field is advancing, the authors point out several "blind spots":

  1. Agile Neglect: Most tools assume heavy documentation (traditional waterfalls). There is a desperate need for tools that work with "User Stories" and "Test-Driven Development" (TDD).
  2. Beyond Text: Modern software includes UI designs and architectural diagrams. TLR needs to evolve to support image-to-text and model-to-code tracing.
  3. Tooling Scarcity: While many prototypes exist (TraceLab, OpenTrace), few offer the "fully automatic" experience promised by ML/DL architectures.

Takeaway

For software architects and researchers, this paper serves as a reminder that traceability is not just a compliance task—it is the neural network of software maintenance. As we move toward AI-assisted coding, integrating these TLR techniques into IDEs will be the key to preventing technical debt from spiraling out of control.

Find Similar Papers

Try Our Examples

  • Find recent papers from 2020-2024 that utilize Deep Learning and Large Language Models (LLMs) to bridge the semantic knowledge gap in software traceability recovery.
  • Which study first introduced the concept of "Transitive Tracing" in software maintenance, and how do modern approaches like CLM (Connecting Link Method) improve upon its original logic?
  • Search for industrial case studies or empirical research that evaluates the effectiveness of automated traceability tools specifically within Agile and DevOps continuous integration pipelines.
Contents
Automatic Traceability: The Bridge Between Software Evolution and Impact Analysis
1. TL;DR
2. The "Knowledge Gap" Problem
3. Methodology: Mapping the Landscape
3.1. The Bohner Model Integration
4. Key Insight: Direct vs. Transitive Tracing
5. Experimental Results & Trends
6. Critical Analysis & Future Outlook
7. Takeaway