Neural Compression & Classification: Solving the Small Data Dilemma in Paralinguistics
Neural Networks for Compressing and Classifying Speaker-Independent Paralinguistic Signals
This paper introduces a neural network-based framework for classifying speaker-independent paralinguistic signals, specifically focusing on heartbeat anomalies and atypical affect detection. It utilizes Autoencoders (AE) for feature compression and Multi-layer Perceptrons (MLP) for classification, achieving SOTA results by effectively handling high-dimensional audio features in small datasets.
TL;DR
Recognizing paralinguistic signals—such as the nuance in a disabled person's voice or the subtle rhythm of a heartbeat—is notoriously difficult because the signs are often "hidden" even from the human ear. This paper addresses the challenge of high-dimensional audio features and small training sets by using Neural Autoencoders for compression and MLPs for classification. The result? Better performance with less data.
Problem & Motivation: The Curse of High Dimensionality
In the world of computational paralinguistics, we often use feature sets like ComParE (6,373 dimensions) and emobase (1,582 dimensions). These sets include everything from prosody and energy to complex cepstrum coefficients.
However, datasets for specific medical or psychological conditions are often tiny (e.g., the Heart Sounds Shenzhen corpus has only 502 training instances). When the number of features is 10x larger than the number of samples, traditional machine learning models like SVMs or Decision Trees tend to memorize noise rather than learning patterns—a classic case of overfitting.
Methodology: Beyond Linear Reduction
The authors compare classic dimensionality reduction techniques like PCA (Principal Component Analysis) and LDA (Linear Discriminant Analysis) against Neural Autoencoders (AE).
1. The Neural Autoencoder Advantage
Unlike PCA, which is a linear transformation, the AE uses non-linear layers (with SELU activation) to "squeeze" the 6,373-dimensional features into a much smaller latent space (e.g., 400 or 200 dimensions). This process forces the network to discard noise and keep only the most vital information for reconstruction.
2. The MLP Classifier
While many practitioners default to XGBoost or SVM for small-scale tabular-style data, this paper argues for the Multi-layer Perceptron (MLP). By using modern techniques like Batch Normalization and Dropout, the MLP can extract high-level representations from the compressed latent vectors that classical models simply miss.

Experiments: Performance Evaluation
The framework was tested on two distinct tasks:
- Heartbeat Classification: Categorizing sounds as "Normal," "Mild," or "Severe" heart disease.
- Atypical Affect Classification: Recognizing emotions (Angry, Happy, Sad, Neutral) in speakers with disabilities.
Key Findings:
- Neural Over Traditional: The MLP almost always outperformed Logistic Regression (LR), Support Vector Machines (SVM), and Gradient Boosting (XGB).
- Compression Gain: Surprisingly, using AE-compressed features often yielded higher accuracy than using the original, full-dimensional features. In the Heartbeat task, AE-compressed features with an MLP reached the top accuracy of 57.67%.
Visualization: Why AE Works Better than PCA
The authors provided a fascinating visual comparison between PCA and AE compression.
Fig. 2. Visualization of AE compression showing cleaner clustering and linear alignment compared to PCA.
As seen in the latent space visualizations, the AE-compressed data points belonging to the same classes were more tightly clustered and better separated than their PCA counterparts. Crucially, the AE distribution for the test set remained close to the training set, which explains the superior generalization.
Critical Insight & Conclusion
The true takeaway here isn't just that neural networks are "better," but that feature redundancy is a major hurdle in audio analysis. By using an Autoencoder as a "pre-filter," we can transform a high-dimensional, noisy problem into a low-dimensional, salient one.
Limitations
While effective, the study relies on fixed feature extraction (openSMILE) before the AE. A modern end-to-end approach (e.g., using a CNN or Transformer directly on the raw waveform) might capture even more information, though it would require significantly more data than what was available in these specific tasks.
Future Outlook
This work paves the way for deploying paralinguistic tools in medical IoT devices or assistive technologies where data collection is expensive, but precision—and the ability to handle small sample sizes—is paramount.
