STCS Lexicon: Rethinking Sentiment Analysis Through Topic-Specific Spectral Clustering
STCS Lexicon: Spectral-Clustering-Based Topic-Specific Chinese Sentiment Lexicon Construction for Social Networks
The paper introduces STCS Lexicon, a novel sentiment lexicon construction method specifically designed for social networks using spectral clustering and a multi-factor similarity graph. It overcomes topic-dependency issues in sentiments and achieves superior performance compared to universal lexicons like NTUSD.
TL;DR
Static sentiment dictionaries are often "blind" to context. The STCS Lexicon framework solves this by constructing topic-specific sentiment lexicons using a sophisticated graph-based approach. By integrating text filtering (FT), multi-dimensional relationship modeling (CRM), and spectral clustering (SC), it accurately classifies words into up to five sentiment intensities, outperforming traditional tools like NTUSD in social network environments.
Background & Motivation: The Context Trap
In linguistic sentiment analysis, the "Holy Grail" has often been a universal dictionary. However, the authors argue that the sentiment of a word is inherently dependent on its topic. For instance, the word "smooth" is highly positive for a smartphone's performance but might be neutral or even negative in a different context.
Current methods suffer from:
- Ambiguity: Words change polarity across domains.
- Inflexibility: Static lexicons cannot keep up with social media slang.
- Coarseness: Most lexicons only offer a binary (positive/negative) or ternary (neutral) split, failing to capture the nuance of human emotion.
Methodology: The Three-Pillar Framework
The researchers propose a pipeline that transforms raw social media noise into a structured, multi-layered sentiment resource.
1. The FT Model (Filtering Text)
Not every comment is useful. The FT model calculates a Text Influence Value () based on user participation (likes, forwards, comments) and persistence. This ensures that the lexicon is built on "hot" and relevant data, filtering out advertisements and noise.
2. The CRM Model (Building the Sentiment Graph)
This is the core "intelligence" of the paper. Instead of relying on a single metric, the authors use a weighted fusion of three factors to calculate word similarity:
- Base Sentiment Similarity: Leveraging existing authoritative lexicons.
- Topic Sentiment Similarity: Calculated by building a "location relationship graph" where the distance between words and the presence of "privative words" (negation) determine similarity.
- Synonym Similarity: Using Jaccard coefficients on synonym sets.

3. The SC Model (Spectral Clustering)
By treating words as nodes and similarities as edge weights, the problem of lexicon construction becomes a graph segmentation problem. Spectral clustering is used to partition the graph into three or five subsets (e.g., Very Positive, Positive, Neutral, Negative, Very Negative).
Experimental Proof: Beyond NTUSD
The authors validated their model using datasets from JD.com ("iPhone6" topic) and eLong ("Hotel" topic).
Key Performance Metrics:
- Precision: In the "Hotel" topic, the method reached 80.1% precision vs. NTUSD's 73.8%.
- Recall: Significant gains were observed, proving that topic-specific modeling captures more relevant sentiment indicators.
- Fine-grained Analysis: The system successfully identified core "Key Words" for different sentiment intensities, such as "satisfactory" vs. "good" or "bad" vs. "not good".

Critical Insight & Future Outlook
The true value of the STCS Lexicon lies in its semi-supervised nature. It reduces the manual labor of labeling while allowing the lexicon to evolve with the conversation. By treating sentiment as a local graph property rather than a global constant, it aligns better with the chaotic reality of social network linguistics.
However, there are limitations. The current model relies heavily on Chinese-specific linguistic features (like "privative word" counts). Future evolution of this work likely involves moving towards Deep Spectral Clustering where the similarity features are learned via embeddings (Word2Vec or BERT) rather than manually defined similarity factors.
Conclusion
By moving from a "One Size Fits All" dictionary to a "Topic-Tailored" graph, Zhang et al. provide a robust blueprint for modern sentiment analysis that respects the complexity of human language in the social media era.
