What Emotions Make One or Five Stars? Unmasking Model Bias through XAI

What Emotions Make One or Five Stars? Understanding Ratings of Online Product Reviews by Sentiment Analysis and XAI

2020-01-01
Chaehan So
Summary
Problem
Method
Results
Takeaways
Abstract

The paper investigates the relationship between emotional sentiments in online reviews and star ratings using machine learning. It employs Explainable AI (XAI) techniques to analyze models like Random Forest and XGBoost, identifying "Joy" and "Negative Valence" as the primary predictors for Amazon product ratings.

TL;DR

This study moves beyond simply predicting star ratings from review text. While Random Forest models can accurately guess a product's rating using sentiment features like "Joy" or "Anger," the author uses Explainable AI (XAI) to prove that these models are often "right for the wrong reasons." The research uncovers that severe class imbalance in Amazon reviews (too many 5-star ratings) leads models to develop flawed logic that can only be caught through local diagnostic tools.

Background: The Black Box of Consumer Sentiment

In the world of e-commerce, reviews are gold. Most research focuses on Accuracy: "Can we predict a 5-star rating based on text?" However, this paper asks a more critical question: "Why did the model think this was 5 stars?" By leveraging NLP to extract emotions (based on Ekman’s universal emotion theory) and emotional valence (positive/negative), the study attempts to bridge the gap between human feeling and machine prediction.

Methodology: Auditing the Machine

The research followed a three-step workflow:

  1. Benchmarking: Testing algorithms like KNN, SVM, Random Forest, and XGBoost.
  2. Global Analysis: Using Feature Importance to see which emotions matter most across the whole dataset.
  3. Local XAI: Using Local Feature Attributions and Partial Dependency Plots (PDP) to see how the model behaves for a single specific review.

Model Benchmarking Results Figure 1: Benchmarking shows Random Forest (rf) achieving the lowest RMSE, but this is only half the story.

The "Aha!" Moment: When Logic Breaks

The Global Feature Importance (Figure 2) suggested that Joy and Negative Valence were the strongest predictors. On the surface, this makes sense. But when the author looked at individual cases using Local Attributions, the model's "thinking" appeared erratic.

Global Feature Importance

For instance, in some cases, a high score for positive valence actually decreased the predicted rating. This is a logical contradiction. The Partial Dependency Plots (Figure 4) further revealed that while "Anger" and "Fear" were correctly linked to lower ratings, the relationships were surprisingly weak, and "Positive Valence" often had a zero slope—meaning the model was effectively ignoring it.

Local Feature Attributions Figure 3: Local attribution reveals that a high intercept (base rating) dominates the prediction, masking the actual influence of sentiments.

The Culprit: Dataset Bias

Why would a high-performing model have such flawed logic? Study 3 provides the answer. By reframing the task as a classification problem, the author found a No-Information Rate of 64.4%.

In plain English: because the vast majority of Amazon reviews are 5-star ratings, a model can achieve ~70% accuracy just by "guessing" 5 stars most of the time. The model wasn't learning the nuances of human emotion; it was learning the statistical imbalance of the dataset.

Critical Insight & Conclusion

This paper serves as a vital warning for data scientists in the NLP and RecSys space:

  • Performance != Understanding: A low RMSE or high Accuracy does not mean your model has "solved" sentiment.
  • XAI as a Diagnostic: Tools like PDP and Local Attributions are not just for "explaining" results to stakeholders; they are essential for debugging and identifying when a model is leaning on dataset bias (shortcuts) rather than features.
  • The Reality of Reviews: Consumer review data is heavily skewed. Without addressing class imbalance, sentiment analysis remains a game of predicting the majority class.

For future work, the study suggests that we must move beyond basic emotion lexicons and utilize XAI to refine feature sets—discarding features with low importance or illogical local variances to build more robust, fair, and interpretable systems.

Find Similar Papers

Try Our Examples

  • Find recent papers that use SHAP or LIME to detect bias in sentiment analysis models for e-commerce.
  • Which study first established the use of Partial Dependency Plots (PDP) for auditing fairness in machine learning?
  • Explore how Explainable AI techniques are being applied to identify and mitigate "label bias" in large-scale imbalanced classification datasets.
Contents
What Emotions Make One or Five Stars? Unmasking Model Bias through XAI
1. TL;DR
2. Background: The Black Box of Consumer Sentiment
3. Methodology: Auditing the Machine
4. The "Aha!" Moment: When Logic Breaks
5. The Culprit: Dataset Bias
6. Critical Insight & Conclusion