M-Lexicon: Bridging the Emotional Gap in Myanmar Social Media Analysis

Word-Emotion Lexicon for Myanmar Language

2019-07-30
Thiri Marlar Swe, Phyu Hninn Myint
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces M-Lexicon, the first dedicated word-emotion lexicon for the Myanmar language. It maps Myanmar words to six basic emotions—happiness, sadness, fear, anger, surprise, and disgust—using a programmatic approach based on Facebook status data and reaction mapping results in a final lexicon of 1,947 emotional words with a 73% F-measure.

TL;DR

Recognizing emotions in text is a cornerstone of modern NLP, yet for the Myanmar language, the lack of a standardized lexicon has been a major roadblock. This paper presents M-Lexicon, a programmatically generated word-emotion association resource built from Facebook data. By combining TF-IDF weighting with social media reactions, the authors created a lexicon covering six basic emotions with an average F-score of 73%.

The Challenge: A Silent Linguistic Frontier

Myanmar (Burmese) is a Sino-Tibetan language spoken by over 30 million people, yet it remains "low-resource" in the digital world. Two main factors make emotion detection difficult:

  1. Structural Complexity: Myanmar script is written without spaces between words, making segmentation (identifying where one word ends and another begins) a non-trivial task.
  2. Resource Scarcity: Unlike English, which has massive lexicons like NRC or EmoLex, Myanmar researchers previously had to rely on inaccurate translations or manual labor.

Methodology: From "Reactions" to "Lexicon"

The authors' insight was to leverage the "wisdom of the crowd" found on Facebook. By collecting statuses and their associated reactions, they created a ground truth for emotional expression.

1. Preprocessing and Segmentation

Since Myanmar text has no spaces, the authors used a two-phase process: Syllable Segmentation followed by Syllable Merging based on a dictionary. Unnecessary "stop words" (particles, pronouns) were removed to focus solely on emotion-carrying terms.

Syllable Segmentation Example

2. The TF-IDF Approach

The system treats word association as a matrix multiplication problem.

  • Word-by-Status Matrix: Calculated using TF-IDF (Term Frequency-Inverse Document Frequency). This identifies which words are unique and significant to specific posts.
  • Reaction Mapping: Facebook reactions were mapped directly to Ekman's six basic emotions (e.g., "Love" and "Haha" Happiness; "Sad" Sadness).

Methodology Flowchart

3. Normalization

To finalize the lexicon, the resulting values were passed through Unity-based (Min-Max) normalization. This scaled the raw scores into a [0, 1] range, effectively allowing the system to set a threshold for whether a word is "associated" (1) or "not associated" (0) with an emotion.

Experimental Performance

The initial M-Lexicon was built using 485 statuses containing 3,890 words, eventually distilling down to 1,947 unique emotional words.

EmotionPrecisionRecallF-measure
Happiness0.8260.8110.818
Anger0.8270.7970.812
Sadness0.7900.7780.784
Fear0.5320.5120.522

The results show high accuracy for "loud" emotions like Happiness and Anger. However, "Fear" and "Disgust" proved more elusive, likely because users express these emotions using more nuanced language or less frequently on public social media.

Critical Insight: Why This Matters

The real value of this paper isn't just the 1,947 words—it's the framework. By providing a programmatic way to turn social media interactions into a structured linguistic resource, the authors have provided a blueprint for other low-resource languages to build their own toolsets without needing massive manual annotation budgets.

Limitations & Future Work

  • Dataset Size: 485 statuses is a small sample. Expanding this to millions of posts would significantly improve the "Fear" and "Disgust" categories.
  • Sarcasm: The current model likely struggles with sarcasm or mixed emotions, a common feature of social media.

Conclusion

M-Lexicon is a vital step forward for Myanmar NLP. It transforms social media noise into structured data, enabling future developers to build better mental health monitoring tools, brand sentiment trackers, and more empathetic AI for Myanmar speakers.

Find Similar Papers

Try Our Examples

  • Search for recent papers on sentiment analysis and emotion detection specifically for the Myanmar language published after 2020.
  • What are the current SOTA methods for Myanmar word segmentation and how do they improve upon rule-based syllable merging?
  • Explore studies that use multi-task learning or cross-lingual embeddings to improve emotion lexicons for low-resource Sino-Tibetan languages.
Contents
M-Lexicon: Bridging the Emotional Gap in Myanmar Social Media Analysis
1. TL;DR
2. The Challenge: A Silent Linguistic Frontier
3. Methodology: From "Reactions" to "Lexicon"
3.1. 1. Preprocessing and Segmentation
3.2. 2. The TF-IDF Approach
3.3. 3. Normalization
4. Experimental Performance
5. Critical Insight: Why This Matters
5.1. Limitations & Future Work
6. Conclusion