Harmonizing Words and Feelings: A Unified SSL Framework for Joint ASR-SER
Semi-Supervised Learning for Multimodal Speech and Emotion Recognition
This paper presents a PhD research plan for a joint Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER) framework using multimodal features (acoustic, visual/lip-reading, and lexical). The core strategy involves a novel cross-task semi-supervised learning (SSL) approach to leverage massive amounts of unlabeled data, aiming to surpass current SOTA performance in emotion-aware interactive systems.
TL;DR
This research tackles the "data desert" in Speech Emotion Recognition (SER) by proposing a joint architecture with Automatic Speech Recognition (ASR). By utilizing multimodal features (audio, lip-reading, and lexical) and a hierarchical fusion strategy, the author aims to use large unlabeled datasets through a novel cross-task Semi-Supervised Learning (SSL) approach. Early results show a significant jump in accuracy (up to ~66%) when using hierarchical attention over standard methods.
The Problem: Small Data, Big Emotions
In the world of AI, ASR is a mature giant, trained on thousands of hours of data like LibriSpeech. SER, however, is a niche orphan. The gold-standard IEMOCAP dataset contains a measly 12 hours of labeled audio. This creates two major issues:
- Lack of Robustness: Models overfit to small sets and fail in the "wild."
- The Multimodal Gap: While we know that what we say (lexical) and how we look (visual) matters as much as how we sound (acoustic), fusing these effectively without massive labels is historically difficult.
The Insight: Cross-Task Synergy
The author's PhD plan rests on a brilliant intuition: ASR and SER should not be silos.
- Lexical Bridge: ASR provides the "what," which is essential context for "how" (emotion).
- Visual Stabilizer: Lip-reading doesn't change based on emotion as much as pitch does, making it a "ground truth" anchor to help ASR remain accurate even when a speaker is shouting or crying.
- Mutual Supervision: In an SSL setting, if the ASR is confident about the text and the SER is confident about the emotion, they can provide high-quality "pseudo-labels" for unlabeled data, creating a virtuous cycle of learning.
Methodology: Hierarchical & Attentional Fusion
The framework moves away from "shallow fusion" (simply concatenating vectors). Instead, it mirrors human auditory processing—moving from low-level frames to high-level semantics.

Core Components:
- Acoustic: Uses wav2vec 2.0 for self-supervised utterance-level representations.
- Visual: Employs spatio-temporal ResNets for character-level lip-reading.
- Lexical: Extracts hidden states from ASR instead of raw text to maintain robustness against transcription errors.
- Fusion: Uses Co-Attention mechanisms to interchange "Key-Value" pairs between modalities, ensuring the model attends to the most emotionally salient parts of the speech.
Experimental Validation
The preliminary results provide strong evidence for the "Hierarchical" hypothesis. By processing lower-level features through CNN-BLSTM layers before fusing them with high-level BERT embeddings, the model achieves superior performance.
| Feature Type | Fusion Approach | Weighted Accuracy |
|---|---|---|
| wav2vec (Audio only) | None | 57.67% |
| BERT (Text only) | None | 50.75% |
| wav2vec + BERT | Hierarchical Shallow Fusion | 66.63% |
| wav2vec + BERT | Hierarchical Attentional Fusion | 62.06% |
(Note: While HSF performed slightly better than HAF in this snapshot, the author notes that attentional mechanisms are more flexible for the planned tri-modal expansion.)

Critical Analysis & Future Outlook
The Good: The transition to self-supervised backbones (wav2vec 2.0) is a major step forward. The emphasis on "uncertainty modeling" in SSL is the right way to handle the noise prevalent in emotional datasets.
The Challenge: Joint training of ASR and SER is notoriously difficult because emotion often degrades ASR performance. The author’s plan to use lip-reading to "bridge" this gap is clever but computationally expensive.
Takeaway: This work signals a shift in affect computing. We are moving away from building "emotion classifiers" and toward building "perceptive agents" that understand language and feelings as a unified signal.
