Beyond Single-Objective Search: Evolving Pareto-Optimal Ensembles for Emotion Recognition
Self-Configuring Ensemble of Neural Network Classifiers for Emotion Recognition in the Intelligent Human-Machine Interaction
The paper introduces a self-configuring multi-objective optimization framework for feature selection and neural network hyper-parameter tuning in audio-visual emotion recognition. It utilizes evolutionary algorithms (SPEA, NSGA-II, VEGA, and the proposed SelfCOMOGA) to evolve Pareto-optimal ensembles of classifiers.
TL;DR
Recognizing human emotions via audio-visual data is notoriously difficult due to high feature dimensionality and the sensitivity of neural networks to hyper-parameters. This paper proposes a multi-objective optimization (MOO) approach that doesn't just look for accuracy, but evolves a "Pareto set" of models that balance performance against complexity. By combining these models into an ensemble using SVM meta-classification, the authors achieved significantly higher robustness and a 7.1% accuracy boost over traditional single-objective tuning.
The "Curse" of Manual Configuration
In Intelligent Human-Machine Interaction (HMI), the goal is to make machines respond to human emotional states. However, the data—consisting of facial video and vocal prosody—is massive. Models are either too complex (overfitting) or too simple (underperforming).
The authors identify two fatal flaws in prior work:
- Feature selection is often treated as a filter, ignoring how specific sets of features interact with the chosen classifier.
- Hyper-parameter tuning is usually a single-track race for accuracy, leading to bloated models that lack generalization.
Methodology: The Pareto Frontier of Neural Networks
Instead of outputting one "best" model, the authors use Multi-Objective Genetic Algorithms (MOGAs) to find a set of models where no single model can be improved in one objective (e.g., accuracy) without degrading another (e.g., simplicity).
1. Dual-Objective Feature Selection
For feature selection, the GA evolves binary vectors representing feature subsets.
- Objective A: Maximize Classification Rate.
- Objective B: Minimize the number of selected features.
2. Self-Configuring Ensembles
To tune the Neural Networks, they introduced SelfCOMOGA, a hybrid algorithm that divides the population into "islands." These islands use different strategies (SPEA, NSGA-II, VEGA) and compete for computational resources based on their ability to find non-dominated solutions.
Table 1: Comparison of Feature Selection methods across different visual and audio descriptors.
Experimental Results: The Power of Multi-Modal Fusion
The researchers tested their approach on the SAVEE database, utilizing diverse features like QLZM (Zernike Moments) and LBP-TOP (Local Binary Patterns on Three Orthogonal Planes).
Key Findings:
- Feature Efficiency: Multi-objective selection was 5.4% more effective than Principal Component Analysis (PCA), proving that keeping "original" features in a smart subset is better than mathematical transformation in this context.
- Ensemble Superiority: Instead of picking the "knee point" of the Pareto front, they combined all non-dominated networks. The SVM Meta-classifier fusion scheme outperformed simple voting, capturing the nuanced strengths of different networks in the ensemble.
Table 2: Performance of various MOGAs and Fusion schemes. Note the dominance of SVM meta-classification in the Audio + Video segment.
Critical Insight & Conclusion
The true value of this work lies in its Self-Configuring nature. In real-world HMI, we cannot afford to manually tune a model every time a new sensor or dataset is introduced. By framing model configuration as a multi-objective competitive evolution, the system automatically adapts its architecture to the difficulty of the task.
Takeaway: If your model is struggling with high-dimensional noise, stop looking for the one "perfect" configuration. Instead, evolve a Pareto-optimal ensemble where different models specialize in different aspects of the feature space.
