Bridging the Semantic Chasm: Recovering Code-to-Doc Traceability with IR
Recovering traceability links between code and documentation
This paper presents an Information Retrieval (IR) based method to automatically recover traceability links between source code and natural language documentation. By utilizing Vector Space and Probabilistic IR models, the authors successfully map C++ and Java classes to manual pages and functional requirements, achieving state-of-the-art results in semi-automated documentation linking.
TL;DR
Maintaining the link between what code is and what documentation says it does is a chronic headache in software engineering. This paper introduces a semi-automated approach using Information Retrieval (IR) models to recover these links. By treating code identifiers as "queries" and documentation as "documents," the authors demonstrate that we can achieve 100% recall in finding relevant documentation with minimal human intervention.
Background: The Concept Assignment Problem
In large-scale legacy systems, documentation is almost always "informal"—written in natural language, free text, or manual pages. When a developer needs to perform impact analysis or bug fixing, finding the specific requirement or manual page associated with a specific Java or C++ class is like finding a needle in a haystack.
The authors' core insight is simple yet powerful: Programmers use meaningful names. Mnemonics like AmountDue or calculate_tax carry domain knowledge. If we can clean these identifiers and match them against the vocabulary of the documentation, we can bridge the semantic gap.
Methodology: The IR Pipeline
The authors propose a dual-path pipeline to normalize both code and documentation into a shared "language."
1. The Normalization Process
Before any matching happens, both artifacts undergo:
- Identifier Splitting: Converting
AmountDueintoamountanddue. - Stop-word Removal: Stripping out "the," "is," "at," etc.
- Morphological Analysis: Stemming words (e.g., "calculating" to "calculate") to ensure they match regardless of grammar.

2. The IR Models
The paper compares two classic IR approaches:
- Vector Space Model (VSM): Uses tf-idf (term frequency-inverse document frequency). Documents and code components are represented as vectors in a high-dimensional space. Similarity is measured by the cosine of the angle between these vectors.
- Probabilistic Model: Uses a stochastic language model (unigram) to estimate the probability that a document is relevant given a specific set of code identifiers.
Experimental Results: Precision vs. Recall
The methodology was tested on two distinct systems:
- LEDA: A C++ library for data structures (95 KLOC).
- Albergate: A Java hotel management system (20 KLOC) mapped to functional requirements.
Key Findings
- The Failure of Grep: Traditional string matching (grep) failed spectacularly. It either found nothing or returned so many results that it was useless for a human maintainer.
- Recall Dominance: Both the Probabilistic and Vector Space models were able to hit 100% Recall (finding every single correct link) by asking the human to review only a handful of top-ranked candidates.

As shown in the table above, with a "Cut" of 12 (reviewing only the top 12 results), the Vector Space model successfully recovered every single link in the LEDA library.
Deep Insight: Why Not Just Use Grep?
One might argue that a simple search is enough. However, the paper proves that IR's ranking mechanism is what provides the value. While grep treats all matches as equal, IR models use global statistics (like how rare a word is across the whole project) to surface the most relevant documents first.
For example, in the Albergate study (performed in Italian), the morphological analysis was crucial. Without converting conjugated verbs to their roots, the "distance" between the code and requirements was too high for any tool to bridge.
Conclusion & Future Outlook
This work establishes a foundational baseline for Software Traceability. It moves beyond syntax and delves into the "informal" semantics hidden in naming conventions.
Takeaways for the Industry:
- Naming Matters: Clean, mnemonic naming is not just for readability; it's for machine-assisted maintenance.
- Effort Saving: Using this IR approach can save over 60% of the effort compared to manual documentation auditing.
Limitations: The model relies on the assumption that code and docs share a vocabulary. If developers use internal jargon that never appears in the requirements, the link breaks. Modern extensions of this work now look at using Synonym Dictionaries and Word Embeddings to solve this specific "vocabulary mismatch" problem.
