PK-Emotion: Bridging Lexical Affinity and Kernel Methods for Chinese Sentence Emotion Recognition

Recognizing sentence emotions based on polynomial kernel method using Ren-CECps

2009-09-01
Changqin Quan, Fuji Ren
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a sentence-level emotion recognition method for Chinese text utilizing the Ren-CECps corpus and a Polynomial Kernel (PK) approach. It classifies sentences into eight basic emotional categories (expect, joy, love, surprise, anxiety, sorrow, angry, hate) by computing similarities between sentences and emotion-specific lexicons, achieving a 62.7% F-measure.

TL;DR

Recognizing specific emotions in text is far more complex than simple sentiment (positive/negative) analysis. This paper presents a robust framework for Chinese sentence-level emotion recognition using the Ren-CECps corpus and a Polynomial Kernel (PK) method. By mapping sentences and emotion-specific lexicons into a high-dimensional feature space, the system identifies eight distinct emotions with a 62.7% F-measure, outperforming standard lexicon-based approaches.

The Challenge: Beyond "Happy" and "Sad"

Most existing sentiment analysis tools focus on document-level polarity. However, understanding the nuance of a sentence—where one might feel both anxiety and sorrow—requires finer granularity. The authors identify two major hurdles:

  1. Indirect Affective Words: Words like "Spring" aren't inherently emotional but carry weight (expect/joy) in specific contexts.
  2. Lexical Coverage: Standard resources like HOWNET are often too formal for the informal, emotion-rich language found in weblogs (blogs).

Methodology: The Power of Polynomial Kernels

The authors propose a kernel-based approach to compute similarity. Instead of simple word matching, they represent documents and emotion lexicons in a Euclidean space where non-linear relationships can be captured.

1. The Ren-CECps Corpus

The foundation is a corpus of 1,487 blog articles (35,096 sentences) annotated for eight emotions. Crucially, they annotate the intensity of emotions (0.0 to 1.0), allowing for the creation of weighted emotion vectors.

2. The Kernel Mathematical Logic

The core similarity is defined by a Polynomial Kernel: This allows the model to capture higher-order interactions between terms (tf-idf) and the emotional weights assigned to those terms.

Architecture of the PK Emotion System Figure 1: The architectural workflow from text preprocessing to the generation of the Polynomial Kernel (PK).

3. Experiential Knowledge

A unique "human-in-the-loop" insight used here is the Emotion Number Distribution. Statistics from the corpus show that most sentences have only 1.36 emotions. The authors use this to set thresholds () to decide whether to assign one, two, or three emotion labels to a sentence, preventing "label explosion."

Experimental Quantitative Results

The model was tested against 9,247 sentences. The comparison between the Ren-CECps lexicon and the widely-used HOWNET is striking:

MetricRen-CECps LexiconHOWNET
Total Words19,0628,936
High-Freq Occurrence (>5)18.4%14.5%
Low-Freq Occurrence (<2)68.2%77.4%

The higher coverage of Ren-CECps in real-world scenarios directly translates to better recognition performance. The system achieves a 62.7% F-measure, which is significant given that human annotators themselves only agree on sentence-level emotion labels about 75.6% of the time.

Emotion Lexicon Examples Table 1: Examples of emotional words and their intensities. Note how "Wait" (望子成龙) carries high "Expect" intensity.

Critical Insight: Why it Works (and Where it Fails)

The PK method succeeds because it doesn't just look for an emotion word; it looks for the topological similarity between the sentence's word distribution and the cumulative distribution of a specific emotion's lexicon.

The Limitation: The model still struggles with linguistic shifts. For example, a sentence containing "Happy" might actually be "Sad" if preceded by "But..." or a negative modifier. Current error analysis shows that ignoring conjunctions (like "but") and rhetorical questions remains the primary source of false positives.

Conclusion

This research underscores that lexicon quality is the bottleneck for affective computing. By building a specialized corpus (Ren-CECps) and applying robust kernel methods, the authors have provided a viable pathway for high-dimensional emotion recognition that goes beyond simple "positive/negative" classification. Future work involving dependency parsing and context-aware embeddings could likely push these results past the 70% mark.

Find Similar Papers

Try Our Examples

  • Search for recent papers that improve sentence-level emotion recognition in Chinese using Transformer-based architectures and Ren-CECps.
  • Which study first introduced the Ren-CECps corpus, and what are its primary annotation differences compared to the WordNet-Affect framework?
  • Explore how polynomial kernel methods have been adapted or replaced by deep learning attention mechanisms for measuring semantic similarity between text and affective lexicons.
Contents
PK-Emotion: Bridging Lexical Affinity and Kernel Methods for Chinese Sentence Emotion Recognition
1. TL;DR
2. The Challenge: Beyond "Happy" and "Sad"
3. Methodology: The Power of Polynomial Kernels
3.1. 1. The Ren-CECps Corpus
3.2. 2. The Kernel Mathematical Logic
3.3. 3. Experiential Knowledge
4. Experimental Quantitative Results
5. Critical Insight: Why it Works (and Where it Fails)
6. Conclusion