Decoding the Table: Fusing Language and Interaction for Role Recognition

Role Recognition for Meeting Participants: an Approach Based on Lexical Information and Social Network Analysis

2010-02-11
Garg, N. P., Favre, Sarah, Salamin, Hugues, Tür, D. Hakkani, Vinciarelli, Alessandro
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a hybrid framework for automatic role recognition of meeting participants by fusing Lexical Analysis and Social Network Analysis (SNA). Tested on the AMI meeting corpus, the combined approach achieves a classification accuracy of approximately 70% in real-world scenarios involving automated speech recognition and speaker diarization.

TL;DR

In any professional meeting, whether you are the Project Manager or the Industrial Designer, your role is defined by what you say and who you talk to. This paper proposes a system that automatically identifies these roles in recorded meetings. By combining lexical n-gram analysis with Social Network Analysis (SNA), the authors achieved a 70% accuracy rate, proving that even "noisy" automatic transcripts can yield deep social insights.

Background: Roles as Social Patterns

Role recognition isn't just about labeling people; it's the key to better meeting summaries, smart indexing, and thematic segmentation. The authors position this work at the intersection of sociolinguistics and graph theory, moving beyond simple "speaker identification" to "functional role understanding."

The Core Problem: Noise and Sparsity

Existing SOTA methods usually fail in two ways:

  1. Lexical Sensitivity: Relying on words is great, but Automatic Speech Recognition (ASR) still has high error rates (often 35-40%), which can garble semantic meaning.
  2. Structural Sparsity: Social Network Analysis usually requires large groups. In a 4-person meeting, the "network" is very small, making it hard to find meaningful patterns.

Methodology: The Dual-Path Framework

The authors propose a parallel architecture to tackle these issues simultaneously:

1. The Lexical Path (The "What")

Using BoosTexter, the system looks at word n-grams (1, 2, and 3-word sequences). A Project Manager might frequently use "budget" or "schedule," while a Designer focuses on "materials." The model treats this as a multi-class categorization task.

2. The Social Path (The "How")

The researchers use Social Affiliation Networks. Instead of just counting who talks, they divide the meeting into time segments (events). If two people consistently speak in the same segments, their "affiliation" is stronger. This is modeled using a Bernoulli distribution to calculate the likelihood of a person's interaction pattern belonging to a specific role.

Model Architecture Figure 1: The dual-path workflow, combining speaker diarization, lexical mapping, and social affiliation modeling.

Experimental Battleground: The AMI Corpus

The system was tested on 45 hours of "remote control design" meetings. There were four roles:

  • Project Manager (PM)
  • Marketing Expert (ME)
  • User Interface Expert (UI)
  • Industrial Designer (ID)

Key Findings

  • Lexical is King: Using words alone was more accurate (67.1%) than social networks alone (43.1%).
  • SNA is the Safety Net: While LEX is better, SNA is more robust. When ASR errors increased, the SNA path's performance dropped much less than the Lexical path.
  • PMs are Predictable: The Project Manager role was the easiest to identify (up to 84% accuracy in the combined model), likely due to their distinct interaction patterns of overseeing all participants.

Performance Results Table 2: Comparison of accuracy across different roles and data sources (Manual vs. Automatic).

Critical Insight & Future Outlook

The beauty of this approach is its identity independence. The system doesn't need to know who you are; it recognizes your role based on current behavior. This makes it highly portable across different organizational contexts.

However, the "small group" limitation remains a hurdle for SNA. The authors suggest that the next frontier is Multimodal Fusion—integrating video data (who is looking at whom) to complement the audio-lexical cues. In the future, your AI assistant might not just record your meeting, but understand exactly who drove the decision-making process.

Takeaway

By fusing "content" (lexicon) with "structure" (SNA), we move closer to machines that understand human social hierarchies as fluently as we do.

Find Similar Papers

Try Our Examples

  • Search for recent studies on role recognition in meetings that utilize deep learning models like BERT or Transformers to improve upon n-gram lexical analysis.
  • What are the state-of-the-art methods for integrating video-based non-verbal cues (e.g., gaze, posture) with audio features for social role detection?
  • Which papers pioneered the use of Social Affiliation Networks for small-group interaction analysis, and how have they been adapted for larger organization meetings?
Contents
Decoding the Table: Fusing Language and Interaction for Role Recognition
1. TL;DR
2. Background: Roles as Social Patterns
3. The Core Problem: Noise and Sparsity
4. Methodology: The Dual-Path Framework
4.1. 1. The Lexical Path (The "What")
4.2. 2. The Social Path (The "How")
5. Experimental Battleground: The AMI Corpus
5.1. Key Findings
6. Critical Insight & Future Outlook
7. Takeaway