Emotion-Based Priority Prediction: Decoding the Human Factor in Bug Triage

Emotion Based Automated Priority Prediction for Bug Reports

2018-01-01
Qasim Umer, Hui Liu, Yasir Sultan
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces an emotion-based automated approach for predicting bug report priority (P1-P5). By integrating sentiment analysis with traditional NLP features, the authors utilize a Support Vector Machine (SVM) classifier to outperform the state-of-the-art DRONE method on large-scale open-source datasets.

TL;DR

Prioritizing bug reports is a critical yet bottlenecked task in software maintenance. This paper presents an automated system that uses Emotion Analysis to "read between the lines" of bug descriptions. By combining sentiment scores from SentiWordNet with textual features and an SVM classifier, the proposed method achieves a 6.1% F1-score improvement over previous state-of-the-art models like DRONE.

The Problem: The Bias of the Reporter

In systems like Bugzilla, the "Priority" field (P1 to P5) is often missing or incorrectly assigned. Reporters have varying technical backgrounds and emotional temperaments. Traditional automated triaging tools often treat bug reports as "cold" technical documents, overlooking the subjective urgency expressed through the reporter's tone. This leads to miscalculations where critical bugs are buried and trivial ones are elevated.

Methodology: The Core Insight

The researchers' key intuition is simple but profound: Severe bugs frustrate users. A report filled with negative sentiment ("bad," "wrong," "suffer") is statistically more likely to be a high-priority issue.

The Workflow

The approach follows a rigorous tripartite pipeline:

  1. Preprocessing: Standardizing text via tokenization, stop-word removal, and Porter’s stemming algorithm.
  2. Emotion Value Calculation: Utilizing SentiWordNet to map words to positive and negative scores. The final "Emotion Value" becomes a quantitative feature alongside traditional bag-of-words tokens.
  3. Classification: Training a Support Vector Machine (SVM). SVM was chosen for its robustness in high-dimensional spaces and its ability to handle the imbalanced nature of bug priority datasets.

Model Architecture Figure: The overall architecture of the emotion-based priority prediction system.

Experiments and Results

The study analyzed 80,000 bug reports from the Eclipse ecosystem (CDT, JDT, PDE, and Platform).

SOTA Comparison

The proposed model was compared against DRONE, the previous benchmark. The results were clear: the inclusion of emotion features led to better precision and recall across the board.

MetricProposed MethodDRONEImprovement
Macro F1-Score46.36%40.26%+6.10%
Micro F1-Score46.31%40.12%+6.19%
Error (Hamming Loss)0.43970.4790-0.082

The Correlation of Sentiment

A key finding was the strong positive correlation (r = 0.405) between emotion and priority. Specifically, 61.71% of P1 bugs contained strong negative sentiment, whereas P5 (lowest priority) reports were dominated by positive or neutral tones.

Experimental Results Figure: Performance comparison across different classification algorithms showing SVM dominance.

Critical Analysis & Conclusion

This paper successfully demonstrates that reporter sentiment is a viable proxy for bug urgency. By quantifying the "human frustration" factor, the model provides a more nuanced triage than models that only look at technical keywords.

Takeaways:

  • Emotional Signals Matter: Text isn't just data; it carries intent. In bug reporting, the way a user complains is as important as what they complain about.
  • SVM Reliability: Even in the era of early deep learning, SVMs remained highly effective for structured feature vectors in software engineering tasks.
  • Limitations: The reliance on SentiWordNet means the model is bound by the quality of the lexicon. Furthermore, sarcasm or highly technical "negative" words (e.g., "kill process") might confound simple sentiment analysis.

Future Outlook: The next logical step is applying Deep Learning (e.g., BERT or LSTMs) to capture the context of these emotion words more effectively, potentially pushing the F1-score even higher.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Deep Learning or Transformers (like BERT) for automated bug report priority and severity classification.
  • What are the historical origins of the DRONE framework in software maintenance, and how has multi-factor analysis evolved since its publication?
  • Are there any studies exploring the application of sentiment analysis to developer comments in issue tracking systems for predicting bug resolution time?
Contents
Emotion-Based Priority Prediction: Decoding the Human Factor in Bug Triage
1. TL;DR
2. The Problem: The Bias of the Reporter
3. Methodology: The Core Insight
3.1. The Workflow
4. Experiments and Results
4.1. SOTA Comparison
4.2. The Correlation of Sentiment
5. Critical Analysis & Conclusion