Failure Triage: Solving the "Bug Infestation" in Modern VLSI Design
A failure triage engine based on error trace signature extraction
The paper introduces an automated "Failure Triage Engine" for semiconductor functional debugging, leveraging SAT-based root cause analysis and hierarchical clustering. It groups large sets of regression test failures by the likelihood of sharing a common root cause, achieving a 93% average classification accuracy.
TL;DR
In the world of high-performance chip design, verification consumes 70% of the development cycle. This paper introduces an automated Failure Triage Engine that uses SAT-based debugging signatures and machine learning (clustering) to group regression failures by their root cause. It achieves 93% accuracy, outperforming traditional script-based sorting by 43%, effectively cutting through the noise of nightly simulation reports.
Background: The Regression Nightmare
In a typical VLSI CAD flow, every code change triggers thousands of nightly "regression tests." When these tests fail, a single logic error might trigger hundreds of different checkers, while multiple independent errors might look deceptively similar.
The industry current relies on two extremes:
- Manual Triage: An experienced engineer reads logs and assigns bugs—accurate but painfully slow.
- Basic Scripts: Grouping failures by error message—fast but notoriously inaccurate (see image below for failure cases).
Figure 1: (a) One error causing different checker failures; (b) Different errors causing the same checker message.
Methodology: Identifying the "Error Signature"
The core innovation is the Failure Proximity metric. Instead of looking at where the error was caught, the engine looks at how it got there.
1. Extracting Propagation Paths
Using modern SAT-based debuggers, the tool identifies "Suspects" (RTL locations) and the "Error Propagation Path." If two failures are caused by the same bug, their propagation paths should converge as we trace them from the chip's outputs back toward the inputs.
2. The Windowing Scheme
The engine divides the error trace into time windows. It calculates the intersection-over-union of suspects in each window.
- Convergence: Ratios increase from outputs to inputs (indicates a shared root cause).
- Divergence: Ratios decrease (indicates distinct sources).
Figure 2: Visualizing how error paths converge toward a single root cause suspect.
3. Hierarchical Clustering (Ward's Method)
The engine uses an agglomerative clustering approach. A key challenge in clustering is knowing the "K" (the number of clusters). The authors developed a heuristic based on Suspect Frequency to estimate the number of actual co-existing errors () without prior knowledge.
Experimental Performance
The framework was tested on designs like a Floating Point Unit (FPU) and a VGA controller.
- Accuracy: Reached 93% using internal signals as suspects.
- Granularity: The tool is sensitive to the number of time windows; 200 windows proved optimal for capturing nuances without introducing noise.
- Efficiency: Total runtime averaged ~15 seconds—a negligible cost compared to hours of manual debugging.
Table 1: Detailed accuracy comparison showing the proposed triage engine significantly outperforming script-based binning.
Critical Insight & Future Outlook
The primary value here lies in the speculative metric of convergence. By moving triage from "textual analysis" (error messages) to "structural analysis" (propagation logic), the authors have turned a heuristic task into a formal one.
Limitations: The algorithm's performance can drop to 86% if the initial "error count estimation" is off. However, since the engine provides a ranked list of suspects, an engineer can quickly spot a misaligned cluster and re-run the process—a "human-in-the-loop" approach that remains far faster than traditional methods.
As chip designs move toward heterogeneous SOCs, this automated triage logic will be vital for managing the sheer scale of functional verification in the AI and 5G eras.
