The Incompleteness Paradox: How Missing Data Shapes Software Traceability
On the effect of incompleteness to check requirement-to-method traces
This paper investigates the impact of data incompleteness on automated requirement-to-method traceability using Machine Learning. By benchmarking Random Forest models on four open-source Java systems, the authors analyze how Undefined (U) traces in training data influence the precision and recall of trace validation.
TL;DR
In the world of Software Engineering, requirement-to-method traces are the "maps" that tell us exactly which line of code implements a specific feature. But what happens when the map is half-empty? This paper explores the impact of incompleteness on Machine Learning models designed to check these traces. Surprisingly, "incomplete" data isn't always a weakness—it can actually help models find more links that a "perfect" dataset would miss.
The Pain Point: The Undefined Reality
Most academic research assumes we live in a binary world: a method either implements a requirement (Trace) or it doesn't (NoTrace). In reality, engineers are busy. They leave thousands of entries Undefined (U).
For a system like iTrust, a developer would need to manually verify over 160,000 potential links. Naturally, they don't finish the job. Previous SOTA methods ignored this "U" state, but this paper argues that the code structure—specifically who calls whom—contains hidden signals that can fill these gaps.
Methodology: Learning from the Call Graph
The authors propose a technique called TraceChecker. Instead of looking at text similarity (which is often misleading), they look at the Method Call Graph.
Core Intuition
If method A calls method B, and method A implements "Payment Processing," there is a high structural probability that B is also involved. The authors extracted 14 features, including:
- Trace values of the owner class.
- Quantities of T, N, and U traces in Callers and Callees.
- Deep neighborhood analysis (Callers of Callers, etc.).

They utilized a Random Forest classifier, which outperformed other models like Naive Bayes and KNN in early testing.
Experiments & The "Complete vs. Incomplete" Battle
The researchers tested their approach on four Java systems: Chess, Gantt, iTrust, and JHotDraw. They created different "Trace Neighborhoods":
- Complete: No "Undefined" traces in the neighborhood.
- Incomplete: Contains one or more "Undefined" traces.
Key Findings
The performance metrics revealed a fascinating trade-off:
| Training Set | Test Set | Precision (T) | Recall (T) | F1-Score |
|---|---|---|---|---|
| Complete | Complete | 87% | 63% | 73% |
| Incomplete | Incomplete | 77% | 71% | 74% |

Why the difference?
- The Precision King: A "Complete" training set is very conservative. It learns strict, clean rules. When it makes a prediction, it's usually right (High Precision), but it's too scared to guess on messy, real-world data (Low Recall).
- The Recall King: An "Incomplete" training set learns how to navigate the "gray areas." It recognizes patterns even when data is missing. This allows it to find significantly more traces, even if it makes a few more mistakes along the way.
Critical Insights & Future Outlook
The most striking takeaway is that incompleteness is a feature, not just a bug. By training on data that looks like the real, messy world, ML models become more robust.
Limitations
- Language Dependency: The current study is focused on Java. While method calls are universal, the parsing tools used (Spoon) are language-specific.
- Call-Graph Only: The model doesn't handle systems where components communicate via message-passing or shared databases, as these don't show up in a standard call graph.
Final Takeaway
If you are building an automated tool to assist developers in software maintenance, don't wait for a "perfect" labeled dataset. Training on the "Undefined" reality of your current codebase might actually result in a tool that discovers more relevant code for your engineers.
