FBP-AL: Elevating Multimodal Emotion Recognition via Factorized Bilinear Pooling and Adversarial Learning

Multimodal Emotion Recognition with Factorized Bilinear Pooling and Adversarial Learning

2021-10-19
Haotian Miao, Yifei Zhang, Daling Wang, Shi Feng
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a novel Multimodal Emotion Recognition framework that combines Factorized Bilinear Pooling (FBP) for sophisticated feature fusion with an Adversarial Learning (AL) mechanism. By integrating ResNet-based visual features and Bi-LSTM-based textual features, the model achieves state-of-the-art performance on the extended FI dataset, specifically targeting 8-category fine-grained emotion classification.

TL;DR

Recognizing human emotion from social media is challenging due to the abstract nature of feelings and the "semantic gap" between raw data and high-level perception. This paper presents FBP-AL, an end-to-end framework that fuses visual and textual data using Factorized Bilinear Pooling and refines these features through Adversarial Learning. The results? A new SOTA on the FI dataset with an 80.7% accuracy.

Context & Motivation

In the era of social networks, emotions are rarely expressed through a single medium. A "happy" caption might accompany a smiling face, or a "beautiful nature" text might complement a serene landscape.

Previous SOTA methods suffered from two main flaws:

  1. Single-modality bias: Relying solely on images or text ignores the complementary nature of multimodal data.
  2. Linear Fusion Limitations: Simply concatenating vectors (Feature Fusion) doesn't allow the model to learn the "math" of how a specific word interacts with a specific image region.

The authors' insight was to treat multimodal fusion as a complex interaction problem and a distribution alignment problem simultaneously.

Methodology: The Core Engine

The FBP-AL architecture consists of two revolutionary components replacing the standard "concatenate and classify" pipeline.

1. Factorized Bilinear Pooling (FBP)

Bilinear pooling is powerful because it calculates the outer product of two vectors, capturing every possible interaction between visual () and textual () features. However, the parameter explosion is massive. FBP solves this by factorizing the weight matrix into low-rank matrices.

Model Architecture Figure: The FBP module facilitates effective inter-modality interaction via sum-pooling and element-wise products.

2. Adversarial Learning (AL) Module

The authors introduce a "hidden" task. Beside the Emotion Classifier (EC), they add a Multimodal Fusion Classifier (MFC).

  • The Game: The MFC tries to identify the "source" of the fusion (e.g., whether the input is a matched pair or a permuted pair).
  • The Goal: By training the generator to "fool" or compete against these classification losses, the model is forced to learn representations that are robust and purely focused on the emotional signal rather than the noise of modality differences.

Experimental Validation

The model was tested on the Extended FI Dataset (22,000+ images with crawled text).

SOTA Comparison

AlgorithmAccuracy (ACC)MAP
Visual (Fine-tuned CNN)58.3%50.4%
Text (SVM)71.0%65.5%
Multimodal (Late-BMA)76.0%69.6%
FBP-AL (Ours)80.7%76.3%

The jump from 76% to 80.7% is substantial in the context of fine-grained (8-class) emotion recognition, where categories like "Excitement" and "Amusement" are notoriously difficult to distinguish.

Ablation Insights

The ablation studies confirmed that using FBP instead of simple concatenation improved performance across almost all categories, notably in "Awe" and "Contentment." Furthermore, the loss curves indicate that adversarial training acts as a "directional guide," leading the model to a deeper local minimum than standard training.

Performance Comparison Figure: Performance comparison across 8 emotional categories showing the superiority of FBP-AL.

Critical Analysis & Conclusion

Takeaway

The strength of this work lies in its move away from "black-box" concatenation toward structured interaction. By using FBP, the model achieves the expressive power of bilinear models without the computational overhead.

Limitations

Despite the gains, the model still struggles with "Anger" and "Fear," often misclassifying them as "Sadness." This suggests that even with FBP, the "affective gap" for high-arousal negative emotions remains a hurdle. Additionally, the textual data was crawled via URLs, which may introduce noise if the metadata isn't strictly emotional.

Future Work

The next step for this lineage of research would likely involve Cross-modal Attention (Transformers) to focus on specific image regions mentioned in the text, further narrowing the semantic gap.

Find Similar Papers

Try Our Examples

  • Find recent papers on multimodal emotion recognition that utilize Transformers or Cross-Attention mechanisms instead of Bilinear Pooling.
  • Which study first introduced Factorized Bilinear Pooling (FBP) for Visual Question Answering, and how has its application evolved in affective computing?
  • Explore newer datasets for fine-grained multimodal emotion recognition that surpass the FI dataset in scale or category diversity.
Contents
FBP-AL: Elevating Multimodal Emotion Recognition via Factorized Bilinear Pooling and Adversarial Learning
1. TL;DR
2. Context & Motivation
3. Methodology: The Core Engine
3.1. 1. Factorized Bilinear Pooling (FBP)
3.2. 2. Adversarial Learning (AL) Module
4. Experimental Validation
4.1. SOTA Comparison
4.2. Ablation Insights
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Work