Decoding the Table: Fusing Language and Interaction for Role Recognition
Role Recognition for Meeting Participants: an Approach Based on Lexical Information and Social Network Analysis
The paper introduces a hybrid framework for automatic role recognition of meeting participants by fusing Lexical Analysis and Social Network Analysis (SNA). Tested on the AMI meeting corpus, the combined approach achieves a classification accuracy of approximately 70% in real-world scenarios involving automated speech recognition and speaker diarization.
TL;DR
In any professional meeting, whether you are the Project Manager or the Industrial Designer, your role is defined by what you say and who you talk to. This paper proposes a system that automatically identifies these roles in recorded meetings. By combining lexical n-gram analysis with Social Network Analysis (SNA), the authors achieved a 70% accuracy rate, proving that even "noisy" automatic transcripts can yield deep social insights.
Background: Roles as Social Patterns
Role recognition isn't just about labeling people; it's the key to better meeting summaries, smart indexing, and thematic segmentation. The authors position this work at the intersection of sociolinguistics and graph theory, moving beyond simple "speaker identification" to "functional role understanding."
The Core Problem: Noise and Sparsity
Existing SOTA methods usually fail in two ways:
- Lexical Sensitivity: Relying on words is great, but Automatic Speech Recognition (ASR) still has high error rates (often 35-40%), which can garble semantic meaning.
- Structural Sparsity: Social Network Analysis usually requires large groups. In a 4-person meeting, the "network" is very small, making it hard to find meaningful patterns.
Methodology: The Dual-Path Framework
The authors propose a parallel architecture to tackle these issues simultaneously:
1. The Lexical Path (The "What")
Using BoosTexter, the system looks at word n-grams (1, 2, and 3-word sequences). A Project Manager might frequently use "budget" or "schedule," while a Designer focuses on "materials." The model treats this as a multi-class categorization task.
2. The Social Path (The "How")
The researchers use Social Affiliation Networks. Instead of just counting who talks, they divide the meeting into time segments (events). If two people consistently speak in the same segments, their "affiliation" is stronger. This is modeled using a Bernoulli distribution to calculate the likelihood of a person's interaction pattern belonging to a specific role.
Figure 1: The dual-path workflow, combining speaker diarization, lexical mapping, and social affiliation modeling.
Experimental Battleground: The AMI Corpus
The system was tested on 45 hours of "remote control design" meetings. There were four roles:
- Project Manager (PM)
- Marketing Expert (ME)
- User Interface Expert (UI)
- Industrial Designer (ID)
Key Findings
- Lexical is King: Using words alone was more accurate (67.1%) than social networks alone (43.1%).
- SNA is the Safety Net: While LEX is better, SNA is more robust. When ASR errors increased, the SNA path's performance dropped much less than the Lexical path.
- PMs are Predictable: The Project Manager role was the easiest to identify (up to 84% accuracy in the combined model), likely due to their distinct interaction patterns of overseeing all participants.
Table 2: Comparison of accuracy across different roles and data sources (Manual vs. Automatic).
Critical Insight & Future Outlook
The beauty of this approach is its identity independence. The system doesn't need to know who you are; it recognizes your role based on current behavior. This makes it highly portable across different organizational contexts.
However, the "small group" limitation remains a hurdle for SNA. The authors suggest that the next frontier is Multimodal Fusion—integrating video data (who is looking at whom) to complement the audio-lexical cues. In the future, your AI assistant might not just record your meeting, but understand exactly who drove the decision-making process.
Takeaway
By fusing "content" (lexicon) with "structure" (SNA), we move closer to machines that understand human social hierarchies as fluently as we do.
