WIYE: Bridging the Generational Gap in Affective Computing through Storytelling
WIYE: building a corpus of children's audio and video recordings with a story-based app
The paper introduces WIYE (What Is Your Emotion?), a multimodal emotional dataset of Italian children aged 4-12. It utilizes a story-based conversational app to collect 1,420 audio/video recordings and 710 emotional utterances across five primary emotions.
TL;DR
Researchers from Politecnico di Milano have developed WIYE (What Is Your Emotion?), the first multimodal emotional corpus comprised of Italian children's recordings. By turning data collection into an interactive "alien-teaching" game, they captured vocal, facial, and semantic data from 142 youngsters, paving the way for more empathetic AI in education and pediatric healthcare.
The Problem: AI's "Adult" Bias in Emotion Recognition
In the realm of Affective Computing, 55% of message perception comes from facial expressions and 38% from tone, yet the vast majority of our training data is derived from adults. This creates a significant "age gap" in model performance. Children express emotions differently—their pitch range, facial muscle movement, and vocabulary are unique. Without age-appropriate corpora, AI systems for schools or therapy remain functionally "blind" to the emotional state of younger users.
Methodology: Gaming the System (For Science)
The researchers recognized that children don't respond well to clinical prompts. Instead, they designed a conversational app starring Boo, an alien who lacks emotions and needs the child's help to learn them.
The Interaction Pipeline
To elicit authentic (though "acted") emotions, the study utilized the Stanislavski method, asking children to recall personal memories associated with joy, sadness, fear, surprise, or anger. Each session involved:
- Imitative Speech: Repeating a context-specific sentence.
- Creative Semantics: Inventing a new sentence based on the current mood.
- Visual Manifestation: Performing a specific facial expression for the camera.
Figure 1: The participant interaction and recording setup for the WIYE corpus.
Experiments & Dataset Insights
The WIYE corpus is characterized by its balanced demographic and high data density:
- Participants: 142 children (71 male, 71 female), aged 4 to 12.
- Quantity: 710 audio recordings for imitative speech, 710 for invented speech, and 710 photos/videos.
- Emotions: Focus on the "Big Six" (minus disgust), aligning with standard industry classifiers like indico.io.
While the researchers acknowledge that "acted" emotions have different signatures than "spontaneous" ones, the inclusion of invented utterances provides a unique semantic layer often missing from simple imitation tasks.
Figure 2: Conceptual representation of the basic emotional states investigated in the story-based app.
Critical Analysis & Future Outlook
The Strength: The method's brilliance lies in its Conversational Agent (CA) design. By positioning the child as the "teacher" to the alien Boo, the power dynamic shifts, increasing engagement and reducing the awkwardness of performing for a machine.
The Limitation: The study notes a risk of over-fitting. Because children invented sentences within a predefined story context, the semantic range is somewhat homogenous (e.g., many space-related sentences).
Future Work: The team plans to expand this into elementary schools to diversify the data and validate the corpus via crowd-sourcing. This research is a cornerstone for creating Emotionally Intelligent Conversational Agents that can assist in diagnosing or supporting children with NeuroDevelopmental Disorders (NDD), where emotional recognition is a vital therapeutic goal.
Conclusion
WIYE is more than just a dataset; it’s a blueprint for ethical, engaging, and age-appropriate data collection. As we move toward a world where AI tutors and healthcare companions are commonplace, datasets like WIYE ensure that the "forgotten demographic" of children is finally heard and understood.
