Decoding Emotions in Marathi: A Statistical Approach to Speech Recognition

Emotion Recognition Using Standard Deviation and Pitch as a Feature in a Marathi Emotional Utterances

2021-01-01
Ashok R. Shinde, Shriram D. Raut, Prashant P. Agnihotri, Prakash B. Khanale
Summary
Problem
Method
Results
Takeaways
Abstract

The paper presents a Speech Emotion Recognition (SER) system specifically designed for the Marathi language using standard deviation and pitch as primary acoustic features. By analyzing 26 specific Marathi utterances across seven basic emotions, the study achieves a high average recognition accuracy of 90%.

TL;DR

Can we identify a person's emotional state just by looking at the standard deviation of their voice? This paper explores this question within the context of Marathi, a regional Indian language. By combining Standard Deviation and Pitch extracted via PRAAT software, the researchers achieved a 90% average accuracy in classifying seven core human emotions, proving that simple statistical markers can be incredibly powerful in localized Human-Computer Interaction (HCI).

Problem & Motivation

While Global languages like English or German have massive emotional speech databases (like Berlin or Spanish databases), regional languages often lack the resources needed for robust AI training. Furthermore, traditional Speech Emotion Recognition (SER) often struggles with:

  1. Complexity: Many models rely on high-dimensional features (like MFCCs) which may capture noise rather than emotional intent.
  2. Overlapping Profiles: Emotions like "Fear" and "Anger" often share similar acoustic energetic profiles, making them hard to distinguish.

The authors hypothesized that the Standard Deviation () of the fundamental frequency—a measure of how much a signal fluctuates from its mean—could serve as a distinct fingerprint for different emotions in Marathi.

Methodology: The Power of Pitch and Deviation

The researchers developed a custom database featuring 50 speakers (25 male, 25 female) recording common Marathi phrases like "Gap re" (Don't talk) and "Are wa" (Great!).

Feature Extraction

The team focused on two key metrics extracted through PRAAT:

  • Standard Deviation (): Used as the primary indicator. Mathematically, it measures the "power" of the fluctuations in the voice signal.
  • Pitch (Fundamental Frequency ): Used as a "tie-breaker." When the standard deviation of two emotions overlapped, the pitch value was used to provide the final classification.

Model Architecture/Workflow Figure: The extraction process using PRAAT software, highlighting the voice portion analyzed for statistical features.

Experiments & Results

The results demonstrated a clear correlation between the statistical variance of a voice and the emotion conveyed.

Key Accuracy Milestones:

  • 100% Accuracy: Happy (25–35 Hz SD), Disgust (37–48 Hz SD), and Surprise (61–80 Hz SD).
  • The Overlap Challenge: The standard deviation for Fear and Angry showed initial similarities. However, by introducing the pitch feature (specifically the 250–257 Hz range for Fear), the system successfully differentiated them.

Emotion Recognition Accuracy Figure: Comparison of recognition rates across the seven studied emotions.

Quantitative Performance Table

EmotionStd. Deviation RangeAccuracy
Happy25–35 Hz100%
Disgust37–48 Hz100%
Surprise61–80 Hz100%
Average-90%

Critical Analysis & Conclusion

The core contribution of this work is the validation of Standard Deviation as a primary feature for SER in the Marathi language. Unlike more complex black-box models, this approach offers interpretability—we can literally see the frequency ranges that define a "Surprise" vs. a "Sad" utterance.

Limitations: The study used a relatively small sample of 26 utterances for the final classification testing. While the results are promising, scaling this to thousands of natural, unscripted conversations remains a challenge.

Future Outlook: This research paves the way for more natural Marathi-speaking virtual assistants. By integrating these statistical boundaries into real-time speech processing, developers can create systems that respond empathetically to the user's tone, specifically tailored to the unique prosody of the Marathi language.

Find Similar Papers

Try Our Examples

  • Find recent papers focusing on Speech Emotion Recognition specifically for Indo-Aryan languages like Marathi or Hindi using deep learning models.
  • What are the established theoretical links between emotional arousal and the standard deviation of fundamental frequency (F0) in prosodic speech analysis?
  • Explore studies that compare the performance of PRAAT-extracted statistical features against end-to-end CNN/RNN architectures in low-resource speech emotion tasks.
Contents
Decoding Emotions in Marathi: A Statistical Approach to Speech Recognition
1. TL;DR
2. Problem & Motivation
3. Methodology: The Power of Pitch and Deviation
3.1. Feature Extraction
4. Experiments & Results
4.1. Key Accuracy Milestones:
4.2. Quantitative Performance Table
5. Critical Analysis & Conclusion