Decoding Internet Slang: A Two-Stage Approach to Chinese Neologism Sentiment Analysis
Associating sentimental orientation of Chinese neologism in social media data
This paper introduces a two-stage framework for identifying Chinese neologisms in social media and determining their sentimental orientation. It combines a statistical discovery method using frequency, user diversity, and duration with a varied TF-IDF algorithm for sentiment association, achieving a 92% recall rate in neologism detection.
TL;DR
Social media is a breeding ground for new linguistic expressions (neologisms) that often bypass traditional dictionaries. This paper presents a specialized pipeline to first discover these words using social behavior statistics (how many people use them over what time) and then associate them with sentimental orientations using a distance-weighted TF-IDF algorithm. The results show a remarkable 92% recall in spotting new words.
The "Moving Target" of Social Media Language
In the landscape of Chinese social media (Sina Weibo, RenRen), language is fluid. Words like "腹黑" (Scheming/Inner Malice) or "坑爹" (To cheat/disappoint) often carry heavy emotional weight but are invisible to standard sentiment lexicons.
The authors identify two core challenges:
- The Discovery Problem: Standard segmenters often break neologisms into meaningless single characters or misidentify them as "unknown words."
- The Sentiment Problem: Even if a word is identified, how do we know if "拉风" (Cool) is positive or negative without a pre-existing label?
Methodology: From Statistics to Sentiment
The researchers proposed a two-step architecture that avoids relying purely on linguistic rules, which are often broken in informal online speech.
Stage 1: Neologism Discovery via Social Behavior
Instead of just looking at how often a word appears (Frequency), the authors introduced the User-Duration Insight. A word is more likely to be a genuine neologism if its use spreads across many different users over a sustained period.
They proposed a scoring formula:
- F (Frequency): Total occurrences.
- U (Users): Number of unique users.
- D (Duration): The lifespan of the word in the dataset.

The top-ranked candidates are then fed into a Support Vector Machine (SVM) that uses the contrast between frequency-based ranking and score-based ranking to confirm the word's status.
Stage 2: Sentimental Orientation Association
To solve the "unknown sentiment" problem, the authors look at the contextual neighborhood. If an unknown word consistently appears near known positive words (like "Happy") and far from negative ones, it inherits that polarity.
The core innovation here is the reciprocal distance weight: The contribution of a known sentimental word to a neologism's score is divided by the distance (offset) between them. If "tragic" is right next to a new word, it has a higher influence than if it's 10 words away.
Experimental Insights
The method was tested on a massive dataset of 3 million microblogs from Sina Weibo.
1. Feature Contribution
The results proved that combined social features () are significantly better than frequency alone. As shown in the comparison, true neologisms (like "Cool" or "Twitter") jumped up the rankings when User Diversity and Duration were factored in.

2. High Recall vs. Low Precision
The SVM classifier achieved a 92% Recall, meaning it successfully caught almost all human-identified neologisms. However, the precision was lower (33%), as the system often flagged celebrity names or unique character combinations (e.g., "地笑") as neologisms.

Critical Analysis & Takeaways
The paper's strongest contribution is the shift from Syntax to Sociology. By treating a word as a "social object" with a life span and a user base, the system bypasses the limitations of rigid grammar rules.
Limitations:
- The accuracy is hampered by the underlying HMM segmenter, which occasionally creates "fragment" neologisms.
- The sentiment precision (50.95%) suggests that "Neutral" words are difficult for the distance-weighting algorithm to distinguish, as they often co-occur with both polarities.
Future Outlook: This framework provides a solid foundation for "Slang Dictionaries" that update in real-time. For practitioners, the key takeaway is clear: if you want to understand social media sentiment, you must track who is using the word and how long it lasts, not just how often it's typed.
