Multi-View Learning: Decoding Emotions in the "Hold 不住" Era of Code-Switching
Multi-view learning for emotion detection in code-switching texts
This paper introduces a multi-view semi-supervised learning framework for emotion detection in code-switching texts (specifically Chinese-English social media posts). By utilizing monolingual views and a synthesized bilingual view via statistical machine translation, the method achieves superior performance in identifying five basic emotions compared to traditional monolingual approaches.
TL;DR
Deep learning for emotion detection often assumes a "pure" linguistic environment. However, real-world social media is a messy blend of languages—Code-Switching. This paper proposes a multi-view learning framework that treats Chinese, English, and a translated "Bilingual" version as distinct perspectives. By leveraging a co-training algorithm, the model learns to bridge the gap between languages, achieving a significant performance boost in identifying emotions like happiness, sadness, and surprise in mixed-language posts.
Problem & Motivation: The Bilingual Emotional Gap
As global social media users frequently mix languages (e.g., "This party was so high 翻全场"), traditional monolingual NLP models face a dilemma. If they focus only on one language, they lose the context of the other; if they treat them as a single string, the statistical sparsity of the secondary language (English in Chinese Weibo) often dilutes the signal.
The authors observed that:
- Monolingual models fail: An English-only model on Weibo data performs poorly (F1-score ~0.32) because English is often used only for specific emotional emphasis.
- Semantic Integration is Key: Emotions in code-switching are not just additive; they are often integrated into specific hybrid phrases that require a unified bilingual understanding.
Methodology: The Three-Lens Perspective
The core innovation lies in the Multi-View Framework, which decomposes a single post into three feature spaces:
- Monolingual Views (CN & EN): Separate features extracted from the native Chinese and English text.
- Bilingual View: This is the "bridge." Since Chinese dominates the dataset, the authors use a Statistical Machine Translation (SMT) strategy to map English words into the Chinese semantic space.
- They don't just stop at translation; they use Sentiment Lexicons and Synonym Dictionaries to ensure that "Happy" (EN) and "开心" (CN) are mapped to the same emotional pivot.
Model Architecture
The framework uses a Co-Training algorithm. Starting with a small set of labeled data, it trains three separate classifiers (). These classifiers then "label" unlabeled data, and the most confident predictions are added back to the training set for the next iteration.

Experiments & Results
The researchers tested their approach on 4,195 manually annotated Weibo posts.
Performance Comparison
The multi-view approach consistently outperformed both supervised baselines and simpler semi-supervised methods.
| Method | Average F1-Measure |
|---|---|
| Baseline (Supervised) | 0.465 |
| ME-CN (Chinese Only) | 0.425 |
| ME-EN (English Only) | 0.325 |
| Multi-View Learning | 0.486 |

Key Insights from results:
- Language Bias: Happiness occurs more frequently in English segments compared to sadness, suggesting a cultural/linguistic preference in how users switch codes.
- The Power of Synergy: Even though the English view alone was weak, it provided complementary "clues" that, when combined with the bilingual view, allowed the model to generalize better than the Chinese-only model (ME-CN).
Critical Analysis & Conclusion
Takeaway: This work proves that in multilingual social media, treating languages as "separate but equal" views is superior to treating them as a single noisy sequence. The use of SMT to create a "middle-ground" bilingual view effectively addresses the data sparsity of the secondary language.
Limitations:
- The translation is word-by-word, which may miss complex idiomatic expressions or slang.
- The framework relies on statistical ME (Maximum Entropy) models, which have since been surpassed by Transformer-based architectures (like mBERT).
Future Outlook: The methodology of "Bilingual Views" is highly applicable to modern Large Language Models (LLMs). By explicitly prompting models to look at code-switching through multiple lingual lenses, we can likely improve the emotional intelligence of AI in globalized digital spaces.
