Failure Triage: Solving the "Bug Infestation" in Modern VLSI Design

A failure triage engine based on error trace signature extraction

2013-07-01
Zissis Poulos, Yu-Shen Yang, Andreas G. Veneris
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces an automated "Failure Triage Engine" for semiconductor functional debugging, leveraging SAT-based root cause analysis and hierarchical clustering. It groups large sets of regression test failures by the likelihood of sharing a common root cause, achieving a 93% average classification accuracy.

TL;DR

In the world of high-performance chip design, verification consumes 70% of the development cycle. This paper introduces an automated Failure Triage Engine that uses SAT-based debugging signatures and machine learning (clustering) to group regression failures by their root cause. It achieves 93% accuracy, outperforming traditional script-based sorting by 43%, effectively cutting through the noise of nightly simulation reports.

Background: The Regression Nightmare

In a typical VLSI CAD flow, every code change triggers thousands of nightly "regression tests." When these tests fail, a single logic error might trigger hundreds of different checkers, while multiple independent errors might look deceptively similar.

The industry current relies on two extremes:

  1. Manual Triage: An experienced engineer reads logs and assigns bugs—accurate but painfully slow.
  2. Basic Scripts: Grouping failures by error message—fast but notoriously inaccurate (see image below for failure cases).

Failure Scenarios Figure 1: (a) One error causing different checker failures; (b) Different errors causing the same checker message.

Methodology: Identifying the "Error Signature"

The core innovation is the Failure Proximity metric. Instead of looking at where the error was caught, the engine looks at how it got there.

1. Extracting Propagation Paths

Using modern SAT-based debuggers, the tool identifies "Suspects" (RTL locations) and the "Error Propagation Path." If two failures are caused by the same bug, their propagation paths should converge as we trace them from the chip's outputs back toward the inputs.

2. The Windowing Scheme

The engine divides the error trace into time windows. It calculates the intersection-over-union of suspects in each window.

  • Convergence: Ratios increase from outputs to inputs (indicates a shared root cause).
  • Divergence: Ratios decrease (indicates distinct sources).

Convergence vs Divergence Figure 2: Visualizing how error paths converge toward a single root cause suspect.

3. Hierarchical Clustering (Ward's Method)

The engine uses an agglomerative clustering approach. A key challenge in clustering is knowing the "K" (the number of clusters). The authors developed a heuristic based on Suspect Frequency to estimate the number of actual co-existing errors () without prior knowledge.

Experimental Performance

The framework was tested on designs like a Floating Point Unit (FPU) and a VGA controller.

  • Accuracy: Reached 93% using internal signals as suspects.
  • Granularity: The tool is sensitive to the number of time windows; 200 windows proved optimal for capturing nuances without introducing noise.
  • Efficiency: Total runtime averaged ~15 seconds—a negligible cost compared to hours of manual debugging.

Results Table Table 1: Detailed accuracy comparison showing the proposed triage engine significantly outperforming script-based binning.

Critical Insight & Future Outlook

The primary value here lies in the speculative metric of convergence. By moving triage from "textual analysis" (error messages) to "structural analysis" (propagation logic), the authors have turned a heuristic task into a formal one.

Limitations: The algorithm's performance can drop to 86% if the initial "error count estimation" is off. However, since the engine provides a ranked list of suspects, an engineer can quickly spot a misaligned cluster and re-run the process—a "human-in-the-loop" approach that remains far faster than traditional methods.

As chip designs move toward heterogeneous SOCs, this automated triage logic will be vital for managing the sheer scale of functional verification in the AI and 5G eras.

Find Similar Papers

Try Our Examples

  • Find recent papers on machine learning techniques applied to automated failure triage in RTL verification beyond hierarchical clustering.
  • Which paper first introduced SAT-based design debugging for RTL, and how does the current work's signature extraction build upon those foundational formal engines?
  • Explore research that applies error trace signature extraction to hardware security verification or side-channel attack diagnosis.
Contents
Failure Triage: Solving the "Bug Infestation" in Modern VLSI Design
1. TL;DR
2. Background: The Regression Nightmare
3. Methodology: Identifying the "Error Signature"
3.1. 1. Extracting Propagation Paths
3.2. 2. The Windowing Scheme
3.3. 3. Hierarchical Clustering (Ward's Method)
4. Experimental Performance
5. Critical Insight & Future Outlook