William the Robot: Crowdsourcing a Coherent and Persistent Personality
Expressing Coherent Personality with Incremental Acquisition of Multimodal Behaviors
The paper introduces an extension to the Persistent Interactive Personality (PIP) architecture, enabling social robots to acquire and express coherent multimodal personalities (Optimistic vs. Impatient) through incremental crowdsourcing. By leveraging a pipeline of crowd workers to author both verbal responses and non-verbal facial expressions, the authors achieved distinct personality traits in a robot named 'William' during a multi-day competitive trivia study.
TL;DR
Researchers have successfully developed a system where a robot’s personality is not hard-coded by a single engineer but "grown" through incremental crowdsourcing. By using a pipeline of crowd workers to author dialog and select facial expressions, the robot "William" displayed two distinct and coherent personas—Optimistic (OPT) and Impatient (IMP)—during a four-day trivia competition, proving that distributed authorship can still result in a unified character.
Backgroud: The Scalability Trap in Social Robotics
In social robotics, personality is the "glue" of long-term interaction. However, creating a personality that feels consistent over weeks or months is a massive content-authoring challenge. Most current systems rely on static, hand-written scripts. When builders try to scale this using crowdsourcing, they often encounter the "incoherence problem": different contributors write in different tones, leading to a robot that feels like it has a multiple personality disorder.
The Problem & Motivation
The core tension lies between Scale and Coherence. How can we collect thousands of interaction lines from strangers on the internet (like Amazon Mechanical Turk) without the robot sounding like a jumbled mess? Furthermore, how do we synchronize these verbal lines with non-verbal cues (facial expressions and voice pitch) to ensure the multimodal signal is clear?
The authors' insight was that personality can be treated as a context variable within a dialog graph, allowing the system to filter and rank contributions that align with a specific psychological narrative.
Methodology: The PIP Architecture
The researchers extended the Persistent Interactive Personality (PIP) architecture. The system functions through a sophisticated loops of learning:
- Semi-situated Learning (The Crowd): Crowd workers are given a "narrative" (e.g., "You are William, an impatient robot who hates wasting time") and asked to write a response to a specific game scenario.
- Embedding Space: Utterances are mapped using doc2vec embeddings. This allows the robot to understand the semantic intent of a user's input and find the most relevant response in its learned graph.
- Multimodal Mapping: Facial expressions weren't just randomly assigned. The robot learned the appropriateness of specific expressions (like a frown or a joyful smile) for every context by having crowd workers rate videos of the robot speaking and moving.
The architecture combines a Dialog Graph with an Embedding Space to manage both social chit-chat and game tasks.
Experiments: Testing OPT vs. IMP
The study involved a 4-day trivia game. William hosted the game, acting either as an optimistic cheerleader or a grumpy, overreacting host.
- The OPT (Optimistic) William: Used faster speech rhythms, phrase-final accents, and joyful facial reactions to "Yes" answers.
- The IMP (Impatient) William: Used a "cross" (angry/annoyed) voice tone, neutral acknowledgments for wins, and impatient reactions to "No" answers.
Results: Distinct and Coherent
The results were striking. Participants didn't just notice a difference; they mapped William onto the "Big Five" personality traits with high precision.
Figure: The IMP personality was rated significantly lower in Agreeableness and Emotional Stability, confirming that the crowd successfully authored a "grumpy" persona.
Key Quantifiable Findings:
- Approval Rate: 98.2% for OPT lines and 96.3% for IMP lines, showing that crowd workers find it easy to "write in character."
- Consistency: Participants rated the match between speech and expression at a Median of 4.0/5.0.
- Engagement: Interestingly, while the IMP robot was rated as "less pleasant," users enjoyed interacting and chatting with it just as much as the OPT version.
Deep Insight: Why Does This Matter?
The real breakthrough here is that distributed authorship does not preclude coherence. By providing narrow narratives and a verification pipeline (Authoring HITs followed by Editing HITs), the system acts as a filter that distills the "wisdom of the crowd" into a singular, predictable character.
However, there are limitations. The personalities tested here (Optimistic vs. Impatient) are quite "loud." It remains to be seen if subtle personality differences—like "slightly introverted" vs "moderately observant"—could be captured so effectively using the same crowd-based approach.
Conclusion
This work provides a blueprint for the future of social robots in our homes and offices. Instead of being locked into a factory-set personality, robots can incrementally acquire new behaviors and traits based on their experiences and human feedback, all while maintaining a coherent "inner self."
