Modeling Speech Emotion: Why Trees Outbranch Regression in Vocal Analysis
Modelling speech emotion recognition using logistic regression and decision trees
2017-09-19
Summary
Problem
Method
Results
Takeaways
Abstract
This paper investigates Speech Emotion Recognition (SER) in the Malayalam language by comparing Decision Tree (CART) and Logistic Regression models. Utilizing vocal tract resonance features—specifically the first four formants and their bandwidths—it achieves a peak accuracy of 93.63% in binary emotion classification.
## TL;DR
This study tackles the challenge of Speech Emotion Recognition (SER) specifically for the Malayalam language. By comparing **Decision Trees** and **Logistic Regression**, the research demonstrates that vocal tract resonances (formants) are better modeled via hierarchical logic than statistical linear links, achieving a stellar **93.63% accuracy** in distinguishing positive emotions.
## Background & Positioning
In the landscape of Human-Machine Interaction (HMI), understanding *how* something is said is as vital as *what* is said. While modern AI often leans toward "black-box" Deep Learning, this paper positions itself in the realm of **interpretable modeling**. It focuses on the Malayalam language (spoken by 35 million people) and seeks to understand the physical contribution of vocal tract changes to emotional expression.
## The Core Motivation: Interpretability vs. Complexity
The author identifies a critical gap: existing SER systems often involve heavy signal processing and high training times (sometimes over 100 hours). The motivation here is two-fold:
1. **Physical Intuition**: Emotional arousal changes respiration and phonation, which directly modifies the shape of the vocal tract.
2. **White-Box Logic**: Can we build a model that doesn't just predict, but tells us *why* it predicted a certain emotion?
## Methodology: Formants as the Fingerprint of Feeling
The researchers extracted the first four **Formants (F1–F4)** and their **Bandwidths (B1–B4)**. Formants are essentially the spectral peaks of the sound spectrum of the voice, representing physical resonances.
### Model 1: Decision Trees (CART)
Using the CART algorithm, the system creates a recursive binary split. It is non-parametric, meaning it makes no assumptions about data distribution—a massive advantage for natural human speech which is often non-linear.
### Model 2: Logistic Regression
This model uses the **Logit Link Function** to map the probability of an emotion into a 0-1 range:
$$\ln \left(\frac{p}{1-p}\right) = \beta_0 + \beta_1x_1 + \dots + \beta_kx_k$$
This provides a strong statistical basis but assumes a specific relationship between predictors and the response.

*Figure 1: The workflow for intuitive modelling of binary emotion classifications.*
## Empirical Results: A Clear Winner
The study conducted experiments on a database of 2,800 speech files from native speakers. The results across three major binary classifications were telling:
* **Neutral vs. Emotional**: Decision Trees achieved ~89% accuracy.
* **Positive vs. Negative Valence**: Decision Trees hit ~91%.
* **Happy vs. Surprise**: This yielded the highest performance at **93.63%**.
In contrast, **Logistic Regression struggled**, often hovering around the 66-73% accuracy range.

*Table 1: Accuracy metrics showing the superiority of Decision Trees over various folds.*
## Critical Insights: Why Did Trees Win?
1. **Non-Linearity**: Human emotion is a complex physiological process. The linear nature of logistic regression (even with logit links) cannot easily capture the "threshold" effects where a specific formant frequency jump suddenly shifts the perceived emotion.
2. **Feature Importance**: Through **Stepwise Regression** and tree pruning, the study found that the **4th Bandwidth (B4)** was often the least significant predictor. This allows for model simplification without losing performance.
3. **Hierarchical Decisions**: Decision trees naturally mimic the way a human might categorize sound—checking one resonance range, then another, creating a path to a leaf node.
## Conclusion & Future Look
The study proves that for specific linguistic contexts like Malayalam, simple spectral features like formants are incredibly powerful when paired with the right architecture.
**Limitations**: The current study focuses on female speech (cited as more expressive) and binary tasks.
**Prospects**: Future research could bridge these "white-box" features with modern Attention mechanisms to see if temporal dynamics improve performance even further without sacrificing the interpretability provided by vocal tract analysis.
