Automatic Traceability: The Bridge Between Software Evolution and Impact Analysis
A Literature Review of Automatic Traceability Links Recovery for Software Change Impact Analysis
This paper presents a Systematic Literature Review (SLR) of 33 primary studies focusing on automatic Traceability Links Recovery (TLR) for Software Change Impact Analysis (CIA). It categorizes methods into Information Retrieval (IR), Heuristic, Machine Learning (ML), and Deep Learning (DL) based approaches, identifying a shift toward DL for bridging knowledge gaps between heterogeneous artifacts.
TL;DR
Software systems are in a constant state of flux. Every change to a requirement or a bug fix can trigger a "ripple effect" across code, design, and tests. This paper provides a comprehensive literature review (2012–2019) of Traceability Links Recovery (TLR) techniques. It identifies how the field is moving from manual keyword matching toward intelligent, deep-learning-based systems that can "understand" the relationship between human language and machine code.
The "Knowledge Gap" Problem
The primary friction in Software Change Impact Analysis (CIA) is the Knowledge Gap. Requirements are written in natural language (informal, ambiguous), while source code is written in programming languages (formal, strict syntax).
Traditional Information Retrieval (IR) approaches assume that if two documents share words, they are related. However, the authors argue this is insufficient for complex evolution where "method names" might not appear in "user stories."
Methodology: Mapping the Landscape
The authors categorized the research into four distinct technological generations:
- Information Retrieval (IR): Using Vector Space Models (VSM) and Latent Semantic Indexing (LSI) to find textual overlaps.
- Heuristics: Adding rules based on change frequency and recency.
- Machine Learning (ML): Training classifiers (Random Forest, SVM) to filter the "noise" (false positives) generated by IR engines.
- Deep Learning (DL): The current frontier, using Word Embeddings and Recurrent Neural Networks (RNNs) to capture deep semantic meanings rather than just keyword counts.
The Bohner Model Integration
A standout feature of this review is how it maps traceability onto Bohner’s CIA Process Model, which divides impacts into four categories:
- RIS (Requirements Impact Set)
- DIS (Design Impact Set)
- TIS (Test Impact Set)
- PIS (Program Impact Set)
The model demonstrates the iterative nature of identifying impacts across different levels of abstraction.
Key Insight: Direct vs. Transitive Tracing
One of the most valuable findings in this review is the distinction in recovery methods:
- Direct Tracing: Works when Source A and Target B are similar.
- Transitive Tracing: Essential when A and B have no common vocabulary. If Requirement A relates to Bug Report C, and Bug Report C relates to Code B, a link is inferred.

Experimental Results & Trends
The review highlights a significant experimental bias:
- 82% of studies use open-source projects.
- Only 9% involve real-world industrial case studies.
- SOTA Achievement: Deep learning approaches are proving superior at handling "polysemy" (words with multiple meanings) and "synonymy," which are the death knells of traditional IR.

Critical Analysis & Future Outlook
While the field is advancing, the authors point out several "blind spots":
- Agile Neglect: Most tools assume heavy documentation (traditional waterfalls). There is a desperate need for tools that work with "User Stories" and "Test-Driven Development" (TDD).
- Beyond Text: Modern software includes UI designs and architectural diagrams. TLR needs to evolve to support image-to-text and model-to-code tracing.
- Tooling Scarcity: While many prototypes exist (TraceLab, OpenTrace), few offer the "fully automatic" experience promised by ML/DL architectures.
Takeaway
For software architects and researchers, this paper serves as a reminder that traceability is not just a compliance task—it is the neural network of software maintenance. As we move toward AI-assisted coding, integrating these TLR techniques into IDEs will be the key to preventing technical debt from spiraling out of control.
