M-Lexicon: Bridging the Emotional Gap in Myanmar Social Media Analysis
Word-Emotion Lexicon for Myanmar Language
The paper introduces M-Lexicon, the first dedicated word-emotion lexicon for the Myanmar language. It maps Myanmar words to six basic emotions—happiness, sadness, fear, anger, surprise, and disgust—using a programmatic approach based on Facebook status data and reaction mapping results in a final lexicon of 1,947 emotional words with a 73% F-measure.
TL;DR
Recognizing emotions in text is a cornerstone of modern NLP, yet for the Myanmar language, the lack of a standardized lexicon has been a major roadblock. This paper presents M-Lexicon, a programmatically generated word-emotion association resource built from Facebook data. By combining TF-IDF weighting with social media reactions, the authors created a lexicon covering six basic emotions with an average F-score of 73%.
The Challenge: A Silent Linguistic Frontier
Myanmar (Burmese) is a Sino-Tibetan language spoken by over 30 million people, yet it remains "low-resource" in the digital world. Two main factors make emotion detection difficult:
- Structural Complexity: Myanmar script is written without spaces between words, making segmentation (identifying where one word ends and another begins) a non-trivial task.
- Resource Scarcity: Unlike English, which has massive lexicons like NRC or EmoLex, Myanmar researchers previously had to rely on inaccurate translations or manual labor.
Methodology: From "Reactions" to "Lexicon"
The authors' insight was to leverage the "wisdom of the crowd" found on Facebook. By collecting statuses and their associated reactions, they created a ground truth for emotional expression.
1. Preprocessing and Segmentation
Since Myanmar text has no spaces, the authors used a two-phase process: Syllable Segmentation followed by Syllable Merging based on a dictionary. Unnecessary "stop words" (particles, pronouns) were removed to focus solely on emotion-carrying terms.

2. The TF-IDF Approach
The system treats word association as a matrix multiplication problem.
- Word-by-Status Matrix: Calculated using TF-IDF (Term Frequency-Inverse Document Frequency). This identifies which words are unique and significant to specific posts.
- Reaction Mapping: Facebook reactions were mapped directly to Ekman's six basic emotions (e.g., "Love" and "Haha" Happiness; "Sad" Sadness).

3. Normalization
To finalize the lexicon, the resulting values were passed through Unity-based (Min-Max) normalization. This scaled the raw scores into a [0, 1] range, effectively allowing the system to set a threshold for whether a word is "associated" (1) or "not associated" (0) with an emotion.
Experimental Performance
The initial M-Lexicon was built using 485 statuses containing 3,890 words, eventually distilling down to 1,947 unique emotional words.
| Emotion | Precision | Recall | F-measure |
|---|---|---|---|
| Happiness | 0.826 | 0.811 | 0.818 |
| Anger | 0.827 | 0.797 | 0.812 |
| Sadness | 0.790 | 0.778 | 0.784 |
| Fear | 0.532 | 0.512 | 0.522 |
The results show high accuracy for "loud" emotions like Happiness and Anger. However, "Fear" and "Disgust" proved more elusive, likely because users express these emotions using more nuanced language or less frequently on public social media.
Critical Insight: Why This Matters
The real value of this paper isn't just the 1,947 words—it's the framework. By providing a programmatic way to turn social media interactions into a structured linguistic resource, the authors have provided a blueprint for other low-resource languages to build their own toolsets without needing massive manual annotation budgets.
Limitations & Future Work
- Dataset Size: 485 statuses is a small sample. Expanding this to millions of posts would significantly improve the "Fear" and "Disgust" categories.
- Sarcasm: The current model likely struggles with sarcasm or mixed emotions, a common feature of social media.
Conclusion
M-Lexicon is a vital step forward for Myanmar NLP. It transforms social media noise into structured data, enabling future developers to build better mental health monitoring tools, brand sentiment trackers, and more empathetic AI for Myanmar speakers.
