Hierarchical Multi-Task Learning: Synchronizing Turn-Level and Call-Level Customer Satisfaction

Customer Satisfaction Estimation in Contact Center Calls Based on a Hierarchical Multi-Task Model

2020-01-01
Atsushi Ando, Ryo Masumura, Hosana Kamiyama, Satoshi Kobashikawa, Yushi Aono, Tomoki Toda
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a Hierarchical Multi-Task (HMT) model for Customer Satisfaction (CS) estimation in contact centers, capable of predicting both turn-level and call-level satisfaction simultaneously. By stacking LSTM-RNNs and using joint optimization, the model achieves SOTA performance on both acted and real-world call datasets.

TL;DR

Analyzing customer satisfaction (CS) in contact centers usually involves two separate tasks: spotting frustrated moments in a specific turn and judging the overall "vibe" of the full call. This paper introduces the Hierarchical Multi-Task (HMT) model, which links these two tasks using LSTMs. By treating call-level satisfaction as a direct consequence of the sequence of turn-level sentiments, the model achieves superior accuracy and proves that we can improve local turn-level detection even when we only have labels for the entire call.

Contextual Positioning

In the world of speech analytics, CS estimation has moved from simple keyword spotting to deep learning. However, most SOTA models treat calls as a "bag of turns" or a single block of audio. This work is a systematic structural improvement that treats a conversation as a hierarchy, mirroring how human supervisors actually evaluate calls.

The Problem: The "Independence" Fallacy

Existing methods typically fail or plateau because:

  1. Context Blindness: They treat each customer turn as an isolated event, ignoring that a customer's anger in Turn 10 is often baked-in starting from Turn 2.
  2. Task Decoupling: Call-level and turn-level predictions are usually trained on separate models. This is counter-intuitive—if a customer ends a call happily, it usually validates the positive turns they just took.

Methodology: The HMT Architecture

The researchers built a stacked network using Long Short-Term Memory (LSTM) units to solve the sequential problem.

1. The Structure

  • Lower Layer (Turn-level): Takes prosodic, lexical, and interactive features and outputs a probability vector for each turn (positive, neutral, negative).
  • Upper Layer (Call-level): Instead of raw audio features, this layer takes the predictions of the turn-level layer as its input.
  • Intuition: The call-level model acts as a "judge" looking at the report card of sentiments across the whole conversation.

HMT Model Architecture

2. Joint Optimization & Adaptation

The model uses a weighted loss function to train both layers simultaneously. The authors also introduced Call-Net Freezing. This is a clever domain adaptation trick: they train a robust model on "acted" (simulated) calls, then freeze the call-level "judge" while only fine-tuning the turn-level "detector" on real-world data.

Joint Optimization Flow

Experiments and Insights

The model was tested on both a high-fidelity acted dataset (29.7 hours) and a challenging real technical support dataset (39.1 hours).

  • The Power of Memory: The LSTM-based approaches consistently crushed traditional SVMs. This confirms that the order of turns matters significantly.
  • Lexical is King: Feature ablation showed that lexical (word-based) features were the strongest indicators of CS, followed by interactive features (like turn-taking speed), while prosody was less reliable in real-world messy audio.
  • The "Trickle-Down" Effect: In a standout experiment, the authors showed that if you only have a label for the whole call (e.g., "The customer was happy"), training the HMT model actually corrects and improves the turn-level detection results.

Performance Comparison Table

Critical Analysis & Takeaways

The HMT model proves that structural Inductive Bias (modeling the turn-to-call hierarchy) is more powerful than just throwing raw data at a generic flat model.

Limitations:

  • The model relies on hand-crafted features (Bag-of-Words and basic prosody), which might miss the nuanced semantic understanding of modern Transformers.
  • Neutral samples heavily dominate the real-world dataset (imbalance issue), which still poses a challenge for high macroF1 scores.

Future Outlook: The logical next step is replacing the feature extraction with an end-to-end transformer (like Wav2Vec 2.0 or HuBERT) while maintaining this hierarchical multi-task head. This would allow the model to catch subtle acoustic cues that current hand-crafted features miss.

Find Similar Papers

Try Our Examples

  • Find recent papers on hierarchical multi-task learning for dialogue sentiment analysis using Transformer-based architectures like BERT or RoBERTa.
  • Which study first proposed the "Joint Many-Task" model in NLP, and how does the HMT model's chained structure differ from that original multi-branch design?
  • Explore research that applies hierarchical sentiment estimation to multi-modal contact center data, specifically combining speech, text, and facial expression analysis.
Contents
Hierarchical Multi-Task Learning: Synchronizing Turn-Level and Call-Level Customer Satisfaction
1. TL;DR
2. Contextual Positioning
3. The Problem: The "Independence" Fallacy
4. Methodology: The HMT Architecture
4.1. 1. The Structure
4.2. 2. Joint Optimization & Adaptation
5. Experiments and Insights
6. Critical Analysis & Takeaways