BUDGET: Automating the Ground Truth for Architecture Traceability Research

BUDGET: A Tool for Supporting Software Architecture Traceability Research

2016-04-01
Joanna C. S. Santos, Mehdi Mirakhorli, Ibrahim Mujhid, Waleed Zogaan
Summary
Problem
Method
Results
Takeaways
Abstract

BUDGET is a web-based tool designed to automate the creation of training datasets for software architecture traceability research. It leverages automated web-mining of technical libraries and big-data analysis of over 22 million source files to identify and extract implementation samples of architectural tactics.

TL;DR

The manual collection of training data is the "Achilles' heel" of automated software traceability. BUDGET (Big-data aUgmented Dataset GEnerator Tool) solves this by mining technical libraries and ultra-large-scale code repositories (22M+ files) to automatically generate high-quality training sets for architectural tactics with over 90% accuracy.

The Bottleneck: Why Traceability Research is Stalled

Software architecture traceability—the ability to link requirements to specific architectural tactics and code—is vital for safety-critical systems and compliance (e.g., HIPAA). However, supervised learning models for this task require massive amounts of labeled data.

Currently, a researcher might spend three months manually vetting code snippets just to train a classifier for five or ten tactics. This "manual bottleneck" prevents the community from scaling research to complex industrial systems and limits the diversity of programming languages and domains studied.

Methodology: Dual-Engine Data Harvesting

BUDGET takes a two-pronged approach to bypass the manual labor of dataset creation:

1. Web-Mining Technical Knowledge

The tool doesn't just look for keywords; it uses textbook definitions of architectural tactics (like Heartbeat or Resource Pooling) to construct intelligent queries via the Google Search API. It targets high-authority technical sites like MSDN and Oracle Documentation to extract "clean" specifications and API usage examples.

2. Big-Data Code Analysis

For implementation samples, BUDGET mines an astronomical repository of 116,609 projects from GitHub, SourceForge, and Apache.

  • Indexing: It uses a custom-built infrastructure that stems and fingerprints 22 million files.
  • Search: It employs a parallelized Vector Space Model (VSM) to calculate cosine similarity between the tactic's textual description and the implementation code.

Overall Architecture

Performance & Accuracy

The authors validated the tool's output against human peer review. The results demonstrate that automated mining is not just faster, but highly reliable:

  • Accuracy: The Big-Data engine achieved 100% accuracy for tactics like Audit, Scheduling, and Authentication.
  • Speed: Searching across the 22-million-file index takes only a few seconds.
  • Diversity: By mining GitHub and Apache, the tool avoids the "silo effect" of single-project datasets.

Accuracy Comparison

Critical Analysis & Conclusion

BUDGET represents a transition toward Evidence-Based Software Engineering. Its ability to generate negative samples—code that looks like a tactic but isn't—is particularly crucial for training robust classifiers that don't produce excessive false positives.

Limitations:

  • While the tool supports 10 hard-coded tactics, user-defined queries require careful keyword selection to maintain high precision.
  • The Web-Mining approach for "Heartbeat" and "Audit" showed lower accuracy (60%), likely due to the commonality of these terms in generic IT documentation.

Looking Ahead: BUDGET paves the way for "Self-Evolving Research." As GitHub grows, BUDGET's knowledge base grows with it. Future iterations could potentially integrate Deep Learning (e.g., CodeBERT) to further refine implementation detection beyond simple VSM similarity.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize automated dataset generation or weak supervision for software architecture traceability or design pattern detection.
  • What are the foundational papers on using the Vector Space Model (VSM) for software artifact traceability, and how does BUDGET optimize this for big-data scales?
  • Explore how big-data analysis tools like BUDGET are being integrated into DevOps pipelines for automated compliance checking and architectural drift prevention.
Contents
BUDGET: Automating the Ground Truth for Architecture Traceability Research
1. TL;DR
2. The Bottleneck: Why Traceability Research is Stalled
3. Methodology: Dual-Engine Data Harvesting
3.1. 1. Web-Mining Technical Knowledge
3.2. 2. Big-Data Code Analysis
4. Performance & Accuracy
5. Critical Analysis & Conclusion