Decoding Internet Slang: A Two-Stage Approach to Chinese Neologism Sentiment Analysis

Associating sentimental orientation of Chinese neologism in social media data

2015-05-01
Lifeng Huang, Xi Liu, Vincent Ng
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a two-stage framework for identifying Chinese neologisms in social media and determining their sentimental orientation. It combines a statistical discovery method using frequency, user diversity, and duration with a varied TF-IDF algorithm for sentiment association, achieving a 92% recall rate in neologism detection.

TL;DR

Social media is a breeding ground for new linguistic expressions (neologisms) that often bypass traditional dictionaries. This paper presents a specialized pipeline to first discover these words using social behavior statistics (how many people use them over what time) and then associate them with sentimental orientations using a distance-weighted TF-IDF algorithm. The results show a remarkable 92% recall in spotting new words.

The "Moving Target" of Social Media Language

In the landscape of Chinese social media (Sina Weibo, RenRen), language is fluid. Words like "腹黑" (Scheming/Inner Malice) or "坑爹" (To cheat/disappoint) often carry heavy emotional weight but are invisible to standard sentiment lexicons.

The authors identify two core challenges:

  1. The Discovery Problem: Standard segmenters often break neologisms into meaningless single characters or misidentify them as "unknown words."
  2. The Sentiment Problem: Even if a word is identified, how do we know if "拉风" (Cool) is positive or negative without a pre-existing label?

Methodology: From Statistics to Sentiment

The researchers proposed a two-step architecture that avoids relying purely on linguistic rules, which are often broken in informal online speech.

Stage 1: Neologism Discovery via Social Behavior

Instead of just looking at how often a word appears (Frequency), the authors introduced the User-Duration Insight. A word is more likely to be a genuine neologism if its use spreads across many different users over a sustained period.

They proposed a scoring formula:

  • F (Frequency): Total occurrences.
  • U (Users): Number of unique users.
  • D (Duration): The lifespan of the word in the dataset.

Discovering Chinese Neologisms - Step 1

The top-ranked candidates are then fed into a Support Vector Machine (SVM) that uses the contrast between frequency-based ranking and score-based ranking to confirm the word's status.

Stage 2: Sentimental Orientation Association

To solve the "unknown sentiment" problem, the authors look at the contextual neighborhood. If an unknown word consistently appears near known positive words (like "Happy") and far from negative ones, it inherits that polarity.

The core innovation here is the reciprocal distance weight: The contribution of a known sentimental word to a neologism's score is divided by the distance (offset) between them. If "tragic" is right next to a new word, it has a higher influence than if it's 10 words away.

Experimental Insights

The method was tested on a massive dataset of 3 million microblogs from Sina Weibo.

1. Feature Contribution

The results proved that combined social features () are significantly better than frequency alone. As shown in the comparison, true neologisms (like "Cool" or "Twitter") jumped up the rankings when User Diversity and Duration were factored in.

Table 2: Contribution of features

2. High Recall vs. Low Precision

The SVM classifier achieved a 92% Recall, meaning it successfully caught almost all human-identified neologisms. However, the precision was lower (33%), as the system often flagged celebrity names or unique character combinations (e.g., "地笑") as neologisms.

Identification Results

Critical Analysis & Takeaways

The paper's strongest contribution is the shift from Syntax to Sociology. By treating a word as a "social object" with a life span and a user base, the system bypasses the limitations of rigid grammar rules.

Limitations:

  • The accuracy is hampered by the underlying HMM segmenter, which occasionally creates "fragment" neologisms.
  • The sentiment precision (50.95%) suggests that "Neutral" words are difficult for the distance-weighting algorithm to distinguish, as they often co-occur with both polarities.

Future Outlook: This framework provides a solid foundation for "Slang Dictionaries" that update in real-time. For practitioners, the key takeaway is clear: if you want to understand social media sentiment, you must track who is using the word and how long it lasts, not just how often it's typed.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize State Space Models or Graph Neural Networks for Chinese neologism discovery in microblogging platforms.
  • Which study first introduced the Metcalf approach for predicting the success of new words, and how has it been adapted for low-resource NLP tasks?
  • Explore how the distance-weighted sentiment association method can be applied to multimodal sentiment analysis involving emojis and text.
Contents
Decoding Internet Slang: A Two-Stage Approach to Chinese Neologism Sentiment Analysis
1. TL;DR
2. The "Moving Target" of Social Media Language
3. Methodology: From Statistics to Sentiment
3.1. Stage 1: Neologism Discovery via Social Behavior
3.2. Stage 2: Sentimental Orientation Association
4. Experimental Insights
4.1. 1. Feature Contribution
4.2. 2. High Recall vs. Low Precision
5. Critical Analysis & Takeaways