HTG: Revolutionizing Music Emotion Classification with Hash Tag Graphs

Automatic music emotion classification using hashtag graph

2019-09-01
Deepti Chaudhary, Niraj Pratap Singh, Sachin Singh
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a novel Hash Tag Graph (HTG) generation method for automatic Music Emotion Recognition (MER). Using a combination of Gabor functions for normalization and multi-dimensional feature extraction, the system classifies Hindi music into six emotional categories, outperforming traditional CNN and SVM models with a SOTA accuracy of 95.7%.

TL;DR

Music is the language of emotion, but teaching machines to "feel" the difference between a sad melody and an angry rhythm has historically been a challenge. This paper presents the Hash Tag Graph (HTG) method—a novel framework that automates music annotation and classification. By leveraging Gabor-filtered features and a correlation-based graph structure, researchers achieved a staggering 95.7% accuracy on Hindi music datasets, leaving traditional Deep Learning (CNN) and Support Vector Machines (SVM) in the dust.

Background: The Semantic Gap in Music

Music Emotion Recognition (MER) typically falls into two traps:

  1. Manual Labor: Annotating thousands of songs with emotions like "disgust" or "surprise" is subjective and slow.
  2. Inefficient Classifiers: Standard models often fail to capture the nuances of pitch, rhythm, and harmony simultaneously.

The authors identify that the "semantic gap"—the distance between raw audio data and human feeling—can be bridged if we use a more structured, hierarchical approach to data labeling.

Methodology: The Core of HTG

The HTG workflow is divided into a rigorous pipeline of normalization, feature extraction, and graph generation.

1. Signal Enhancement via Gabor Functions

Before features are extracted, the music signal undergoes Normalization and filtering using a Gabor function. This process removes noise and highlights the time-frequency energy distributions critical for emotional perception.

2. Multi-Dimensional Feature Extraction

The system extracts 36 distinct features, including:

  • Timbral Features: Spectral Centroid, Flux, and Roll-off.
  • Energy Distribution: Sub-band power and Zero-crossing rates.
  • Rhythmic Features: Pitch and beat detection (MFCC).

3. The Hash Tag Graph (HTG) Generation

This is the architectural crown jewel. Instead of a "flat" classification, the system builds a tree:

  • Header Node: Based on the maximum correlation value.
  • Positive Branch (+ive): If the threshold exceeds the header node, emotions like Happy and Surprise are categorized here.
  • Negative Branch (-ive): Lower correlation leads to Angry, Disgust, Fear, and Sad.

HTG Training Process Figure 1: The proposed HTG training workflow, showing the path from music database to the final Graph.

Experiments and Results

The study compared the HTG method against three heavyweights: CNN, SVM, and KNN.

SOTA Performance

The superiority of the graph-based approach is evident across every metric:

  • Accuracy: HTG (95.7%) vs. CNN (79.3%).
  • Speed: HTG achieved the lowest computational cost (559.7 seconds), proving that smarter architecture beats raw brute-force computation.
  • Error Rate: HTG maintained an RMSE of 0.56, nearly 3x lower than KNN.

Performance Comparison Figure 2: Statistical comparison across Specificity, Precision, Recall, and F-measure.

Why does HTG work?

In the paper's Ablation-style insights, the authors show that the correlation-based annotation removes the "uncertainty" in the data. By pre-sorting emotions into positive/negative clusters before fine-grained classification, the model avoids the common confusion between high-arousal emotions (like Anger vs. Surprise).

Critical Analysis & Conclusion

Takeaway

The HTG method proves that structural priors (like the positive/negative emotional split) combined with mathematically rigorous filtering (Gabor functions) can outperform generalized Deep Learning architectures in specialized domains like Music Emotion Recognition.

Limitations & Future Work

While the results on the 1,000-song Hindi dataset are impressive, the paper primarily focuses on 30-second clips. Future research should explore:

  • Cross-Cultural Validation: Does an HTG trained on Hindi "Ghazals" work on Western "Jazz"?
  • Temporal Dynamics: How do emotions change within a song (e.g., a crescendo moving from Sad to Angry)?

Ultimately, this work paves the way for smarter music streaming services that can curate playlists not just by "genre," but by the true emotional resonance of the sound.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize graph-based structures or Gabor filters for music emotion recognition in non-Western musical contexts.
  • Which study first introduced the use of correlation coefficients for automatic music annotation, and how does the HTG method specifically iterate upon that foundation?
  • Explore how the Hash Tag Graph approach could be adapted for real-time multimodal emotion detection combining audio features with facial expression analysis.
Contents
HTG: Revolutionizing Music Emotion Classification with Hash Tag Graphs
1. TL;DR
2. Background: The Semantic Gap in Music
3. Methodology: The Core of HTG
3.1. 1. Signal Enhancement via Gabor Functions
3.2. 2. Multi-Dimensional Feature Extraction
3.3. 3. The Hash Tag Graph (HTG) Generation
4. Experiments and Results
4.1. SOTA Performance
4.2. Why does HTG work?
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations & Future Work