Bridging Text and Expression: Building a Japanese Emotion Ontology from the Social Web

Building of Japanese Emotion Ontology from Knowledge on the Web for Realistic Interactive CG Characters

2013-07-01
Kosuke Kaneko, Yoshihiro Okada
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a novel framework for building a Japanese Emotion Ontology by mining knowledge from web sources like Twitter and BBS. It utilizes a Naïve Bayes approach to calculate emotional intensities and represents the data using OWL and EmotionML to drive realistic facial animations for CG characters.

TL;DR

In the quest for realistic interactive CG characters, the ability to "feel" and express emotions based on conversation is paramount. This paper presents a methodology to build a Japanese Emotion Ontology by mining over a million entries from Twitter and BBS (2-Channel). By calculating emotional intensities using weighted probability and structuring them via OWL and EmotionML, the authors provide a bridge that transforms raw Japanese text into nuanced 3D facial animations.

Background: The Role of Emotion in Interaction

As digital concierges and game characters become more integrated into our lives, the "Uncanny Valley" often looms large—not just because of visual fidelity, but because of emotional misalignment. Most existing systems treat emotions as binary (Happy or Sad), but human expression is a spectrum of intensity.

The authors identify a critical gap: the lack of a structured, intensity-aware emotion dataset for the Japanese language. Their solution? Leveraging the "wisdom—and emotion—of the crowds" found on social media.

Methodology: Quantifying the Ineffable

The core innovation lies in how the researchers quantified the "strength" of an emotion associated with a word.

1. The Triple-Layer Emotion Model

Instead of choosing one theory, the authors used three to ensure the ontology's versatility:

  • Discrete Categories: 10 distinct emotions (Joy, Anger, Sadness, Fear, Shame, Like, Disgust, Exciting, Comforted, Surprise).
  • Simple Polarity: Positive, Negative, and Neutral.
  • Dimensional Model (PAD): Pleasure, Arousal, and Dominance.

2. Intensity through Probability

The authors didn't just label words; they calculated a Weighted Conditional Probability. By analyzing the occurrence frequency of words within manually tagged emotional documents, they derived a score representating how strongly a word (like Tanoshii - Joyful) correlates to a specific emotion.

Workflow Overview Figure 1: The research pipeline from web-crawling to ontology building and application.

Architecture: From Data to Ontological Knowledge

Using OWL (Web Ontology Language), the researchers ensured that their Japanese dataset could be "linked" or merged with other global ontologies (e.g., English or Chinese datasets). Within these OWL classes, they embedded EmotionML, a W3C standard, to store the calculated intensity values as machine-readable attributes.

Calculating Variation for Animation

The bridge to the physical world (or the digital-physical world) is the MPEG-4 Specification. The system maps the highest-intensity emotion from a sentence to 68 Face Animation Parameters (FAPs).

Facial Animation Results Figure 2: Example facial animations: Neutral, Joy, Sadness, and Anger generated from ontology parameters.

Experimental Results

The authors extracted 1,033,511 texts from Twitter and 2-Channel. By focusing on adjectives—the primary carriers of emotional weight in Japanese—and applying morphological analysis (CaboCha), they created a robust intensity matrix.

A sample result for the word "Joyful" (Raku/Tanoshii) showed a high JOY intensity (0.183), while a neutral sentence like "Today was a joyful day" yielded a lower consolidated probability across multiple categories, demonstrating the system's ability to handle context.

Critical Analysis & Future Directions

Takeaway

This work is a significant step toward semantic emotional interoperability. By using OWL, the authors aren't just building a database; they are creating a component of the "Emotional Semantic Web" where different AI agents can share a common understanding of human feelings across languages.

Limitations

  1. Sarcasm & Context: The current Naïve Bayes approach treats words largely in isolation or simple combinations, which may struggle with the deep sarcasm common on platforms like 2-Channel.
  2. Morphological Focus: By focusing primarily on adjectives, the system might miss nuanced emotional cues found in Japanese verbs or sentence-ending particles.

The Path Forward

The authors propose expanding this to include time, location, and personal context, moving toward a "Linked Life Data" approach. For CG characters, the next frontier isn't just a moving face, but a "moving body" (body motion generation) that matches the ontological state of the character.

Conclusion: This paper proves that the chaotic data of the web is a goldmine for building structured knowledge that makes our digital companions more human.

Find Similar Papers

Try Our Examples

  • Find recent papers that extend the EmotionML and OWL framework for multi-modal emotion recognition in social robots.
  • Which study first introduced the Japanese Emotion Expression Dictionary, and how have recent deep learning models improved its classification accuracy compared to Naïve Bayes?
  • Explore how the PAD (Pleasure-Arousal-Dominance) model is currently being applied to procedural body motion generation in digital humans.
Contents
Bridging Text and Expression: Building a Japanese Emotion Ontology from the Social Web
1. TL;DR
2. Background: The Role of Emotion in Interaction
3. Methodology: Quantifying the Ineffable
3.1. 1. The Triple-Layer Emotion Model
3.2. 2. Intensity through Probability
4. Architecture: From Data to Ontological Knowledge
4.1. Calculating Variation for Animation
5. Experimental Results
6. Critical Analysis & Future Directions
6.1. Takeaway
6.2. Limitations
6.3. The Path Forward