Decoding Emotions in Fossils: A Knowledge-Based Approach to Chinese Idiom Classification
Emotional Classification of Chinese Idioms Based on Chinese Idiom Knowledge Base
The paper introduces the development of the Chinese Idiom Knowledge Base (CIKB) and presents an automatic emotion classification system for idioms. Using a Support Vector Machine (SVM) and a rich feature set, the model classifies idioms into Appreciative, Derogatory, and Neutral categories, achieving a peak F-score of 75.93%.
TL;DR
Researchers from Peking University have developed a systematic approach to classify the emotional polarity of Chinese idioms—categorizing them as Appreciative, Derogatory, or Neutral. By leveraging the Chinese Idiom Knowledge Base (CIKB) and a Support Vector Machine (SVM) classifier, they achieved an F-score of 75.93%. The study reveals that the key to understanding an idiom's emotion lies not in modern word segmentation, but in a hybrid strategy that respects the idiom's ancient character-based roots while exploiting modern textual explanations.
The "Fossil" Problem: Why Idioms Defy Standard NLP
In most Natural Language Processing (NLP) tasks, we assume a degree of compositionality—the meaning of a sentence is the sum of its words. However, idioms are "fossils" of language. Their meanings are figurative and cultural, often preserved from ancient Chinese.
The authors identify two primary challenges:
- Semantic Opacity: The literal meaning of "hanging a sheep's head while selling dog meat" (挂羊头卖狗肉) has nothing to do with butchery; it's a derogatory term for deception.
- Structural Rigidity: Modern segmenters (like ICTCLAS) often fail on idioms because the internal grammar of an idiom follows ancient rules, not modern ones. Applying standard POS (Part-of-Speech) tagging actually introduces noise rather than clearing it up.
Methodology: Bridging the Ancient and the Modern
To solve the classification problem, the researchers utilized a massive dataset of 20,000 idioms for training. The technical core of their approach is a heterogeneous feature engineering strategy:
1. The Classifier
The team used LIBLINEAR (L2-loss SVM), a robust choice for high-dimensional text classification.
2. Feature Hierarchy
They tested three types of features:
- Idiom Characters (i_cu, i_cb): Unigrams and bigrams of the characters within the idiom.
- Explanation Words (e_wu, e_wb): Since idioms are hard to parse, they used the modern Chinese explanation of the idiom as a feature pool.
- POS Tags: Grammatical categories of the constituents.

The design intuition here is brilliant: treat the idiom itself as a sequence of symbols (character-based) but treat its dictionary definition as a modern semantic vehicle (word-based).
Experimental Insights: What Actually Works?
The results provided several counter-intuitive but linguistically sound insights:
- Segmentation Hurts Idioms: Using word-level features (
i_wu) for the idiom itself performed worse than character-level features (i_cu). This confirms that idioms are "frozen" units. - Explanations are Essential: The performance jumped significantly when word features from the explanation field were added. The explanation provides the emotional "clue" that the four-character idiom hides.
- POS Tags are Noisy: Adding POS features decreased the F-score. The authors suggest this is because the archaic grammar of idioms confuses modern POS taggers.

As shown in the table above, the combination i_cu + i_cb + e_wu + e_wb (Idiom characters + Explanation words) yielded the state-of-the-art result for this specific framework.
The Learning Curve
The research concludes with a promising outlook. The learning curve shows that performance peaks at 20,000 idioms but still shows a slight upward trend. This suggests that expanding the CIKB further could lead to even higher accuracy.

Conclusion & Critical Analysis
This work highlights a critical lesson for modern AI: Knowledge matters. While modern LLMs often "guess" sentiment from context, this paper shows that for culturally dense units like idioms, a structured Knowledge Base (CIKB) provides a ground truth that raw statistical parsing cannot reach.
Limitations: The study relies on explicit explanations being available. In a real-world "wild" text scenario, an NLP system might encounter an idiom without a dictionary gloss nearby. The next frontier for this research—which the authors hint at—is Event Classification: understanding not just if an idiom is "good" or "bad," but in what specific social scenarios it is deployed.
Takeaway for Practitioners: When dealing with domain-specific or culturally-fossilized language, prioritize character-level modeling and cross-reference with external knowledge bases rather than relying on modern syntactic parsers.
