Streamlining the Rhythm: Pruning Neural Networks for Music Emotion Recognition
Pruning Long Short Term Memory Networks and Convolutional Neural Networks for Music Emotion Recognition
This paper investigates Music Emotion Recognition (MER) by classifying music genres from physiological signals (electrodermal activity) using Long Short Term Memory (LSTM) and Convolutional Neural Networks (CNN). The authors achieve 69.23% and 72.97% accuracy respectively while demonstrating that model complexity can be significantly reduced through weight pruning without sacrificing performance.
TL;DR
Can your skin's electrical conductance reveal if you are listening to Mozart or Maroon 5? This paper proves it can, using LSTMs and CNNs to classify music genres from electrodermal activity (EDA). More importantly, it reveals that these models are often "overbuilt"—specifically, the LSTM model could be pruned by nearly 98% while maintaining its classification accuracy, paving the way for efficient, lightweight music therapy AI.
Problem & Motivation: The Complexity Trap
In the realm of deep learning, there is a persistent "Goldilocks" problem: how many hidden neurons are just right? For sequence data like physiological signals, researchers often rely on "rules of thumb"—such as setting hidden layers to two-thirds the size of the input plus output.
However, in Music Emotion Recognition (MER), where datasets are often small and signals are noisy (like the 24-participant EDA dataset used here), large models tend to overfit. This study seeks to bridge the gap between high-performance sequence classification and model efficiency by applying Network Pruning.
Methodology: EDA to Genre Classification
The researchers utilized a dataset of electrodermal activity (sweat gland response) collected at 4 Hz. They compared two heavyweights of sequence processing:
- LSTM (Long Short-Term Memory): Chosen for its ability to capture long-term dependencies in 4-minute songs.
- CNN (Convolutional Neural Network): Specifically a 1D-CNN designed to extract local temporal features from the EDA stream.
Architecture and Pruning Strategy
The core of the methodology lies in the pruning of the fully connected layer. For a model with 100 hidden neurons and 3 output classes (Classical, Instrumental, Pop), there are 300 weights. The authors systematically removed these weights to observe the "breaking point" of the model.
Figure: The LSTM (left) reaches peak accuracy early, while the CNN (right) shows a steadier upward trend but requires more epochs.
Experiments & Results: The "Sparse" Truth
The results were striking. The CNN outperformed the LSTM slightly in raw accuracy (72.97% vs 69.23%), likely because spatial/local filters in the CNN effectively filtered out noise in the EDA signal.
The Pruning Breakthrough
The most significant finding came from the pruning experiments.
- LSTM Resiliency: The LSTM accuracy didn't drop significantly until 295 out of 300 weights were removed. This indicates an extreme level of redundancy; only a few "critical neurons" were actually doing the heavy lifting for genre recognition.
- CNN Resiliency: The CNN was less resilient than the LSTM but still allowed for a 20% reduction in structure without performance loss.
Figure: Comparison of accuracy retention as weights are pruned. Note the sharp drop-off only at extreme sparsity levels for the LSTM.
Computational Efficiency
Pruning wasn't just about size; it was about speed. The LSTM showed a clear trend where pruning highly correlated with reduced training time per epoch (up to 17% faster). For the CNN, the benefits were less pronounced in the final layer, suggesting that the bulk of CNN computation resides in the earlier convolutional layers.
Critical Analysis & Conclusion
Takeaways
This research provides a strong Inductive Bias for future physiological signal processing: simpler is often better. The fact that an LSTM can function on <5% of its weights suggests that the underlying "features" of music emotion in EDA data are low-dimensional.
Limitations
- Small Dataset: With only 24 participants, the models are prone to noise, as seen in the "jagged" accuracy curves.
- Truncation: The authors truncated songs to the shortest length. This might discard valuable emotional "crescendos" occurring at the end of longer tracks.
- Feature Breadth: Using only EDA limits the model. As the authors noted, previous work using 16 distinct features achieved 96% accuracy.
Future Outlook
The path forward lies in Multi-modal Pruned Networks. By combining EDA with Heart Rate Variability (HRV) and using pruning techniques, we can develop wearable devices that monitor emotional states in real-time with minimal battery consumption—a holy grail for personalized music therapy.
