Leveraging Multilingual BERT for Political Sentiment in Code-Mixed Tamil Tweets

Exploiting Multilingual Neural Linguistic Representation for Sentiment Classification of Political Tweets in Code-mix Language

2021-06-29
Rajkumar Kannan, Sridhar Swaminathan, Chutiporn Anutariya, Vaishnavi Saravanan
Summary
Problem
Method
Results
Takeaways
Abstract

This paper proposes a sentiment classification framework for political tweets written in "Code-mix" (Tamil-English) and pure Tamil. It leverages a fine-tuned Multilingual BERT (mBERT) model to achieve state-of-the-art accuracy in identifying positive, negative, and neutral sentiments.

    ## TL;DR
    Social media discourse in multilingual societies often happens in "Code-mix"—a hybrid of local languages and English. This paper presents a robust sentiment analysis pipeline specifically for **Tamil-English Code-mix** political tweets. By fine-tuning **Multilingual BERT (mBERT)**, the researchers achieved a **91% testing accuracy**, significantly outperforming traditional machine learning models like SVM or Random Forest.

    ## The Challenge of Code-Mix Social Media
    Analyzing sentiment on Twitter is notoriously difficult due to:
    *   **Linguistic Hybridity**: Users frequently switch between Tamil and English scripts or use English phonetics to write Tamil words.
    *   **Morphological Complexity**: Tamil is a morphologically rich language, making standard stemming ineffective.
    *   **Noise**: Political tweets are cluttered with hashtags, user mentions (@handles), and URLs that dilute sentiment signals.

    The authors argue that existing systems relying on Bag-of-Words or simple lemmatization cannot capture the nuanced context of political opinions expressed in these mixed formats.

    ## Methodology: The mBERT Pipeline
    The researchers developed a multi-stage approach to transform raw, noisy tweets into structured sentiment insights.

    ### 1. Preprocessing & Transliteration
    Unlike English-only pipelines, this method includes:
    *   **Morphological Analysis**: To handle Tamil’s complex word structures.
    *   **iTrans Transliteration**: Converting Tamil characters into a standardized Romanized format to improve the phonetic alignment with English tokens.
    *   **Cleaning**: Removal of non-sentimental noise (URLs, @handles).

    ### 2. Architecture: Neural Linguistic Representation
    The core of the system is the **mBERT (Multilingual Bidirectional Encoder Representations from Transformers)**. Unlike traditional models that look at words in isolation, mBERT looks at the entire context of a sentence simultaneously.

    ![The Proposed Architecture](https://cdn.atominnolab.com/wisdoc/images/20260527-3c521a0f-1e04-4e05-8461-5f0c9f00f3b3/page_002_block_002.png)
    *Figure 1: The system architecture showing the flow from data collection to mBERT feature representation and final classification.*

    The model was fine-tuned on a manually labeled dataset of **16,424 tweets**, allowing the transformer to adapt its pre-trained 104-language knowledge to the specific nuances of Indian political discourse.

    ## Performance Benchmarks
    The study compared the Deep Learning Transformer against various traditional Machine Learning (ML) algorithms.

    | Model Name | Precision (Test) | Recall (Test) | Accuracy (Test) |
    | :--- | :--- | :--- | :--- |
    | **Deep Learning Transformer** | **92** | **89** | **91** |
    | Random Forest | 90 | 83 | 83 |
    | Decision Tree | 73 | 73 | 73 |
    | Support Vector (SVC) | 89 | 78 | 78 |
    | Naïve Bayes | 71 | 70 | 70 |

    ### Key Findings:
    *   **Overfitting in ML**: Random Forest and Decision Trees achieved 100% accuracy on training data but collapsed during testing, indicating they "memorized" the data rather than learning the language patterns.
    *   **Superior Generalization**: The Transformer model maintained high stability between training (94%) and testing (91%), proving its ability to handle unseen code-mixed expressions.

    ## Deep Insight: Why Does mBERT Win?
    The success of mBERT in this context stems from its **shared embedding space**. Because mBERT was trained on 104 languages simultaneously, it learns "cross-lingual" features. When a user writes a Tamil sentiment using English script, mBERT can leverage its understanding of both Tamil semantics and English syntax to bridge the gap—something a standard dictionary-based model could never do.

    ## Conclusion & Future Outlook
    This research provides a vital blueprint for governmental and social organizations to measure public mood in linguistically diverse regions. By moving beyond simple keyword matching to neural linguistic representations, we can finally decode the complex "digital dialect" of the modern web. 

    Future work aims to expand this to other social platforms like Facebook and explore even more advanced architectures like XLM-R to further refine the detection of subtle cultural nuances.

Find Similar Papers

Try Our Examples

  • Find recent papers addressing sentiment analysis in other Dravidian code-mixed languages like Telugu-English or Malayalam-English using XLM-RoBERTa.
  • Which paper first introduced the mBERT (Multilingual BERT) architecture, and how does its training objective differ from the standard monolingual BERT?
  • Explore research that applies transliteration-based preprocessing to improve the performance of Transformers in South Asian code-mixed text classification.
Contents
Leveraging Multilingual BERT for Political Sentiment in Code-Mixed Tamil Tweets
1. TL;DR
2. The Challenge of Code-Mix Social Media
3. Methodology: The mBERT Pipeline
3.1. 1. Preprocessing & Transliteration
3.2. 2. Architecture: Neural Linguistic Representation
4. Performance Benchmarks
4.1. Key Findings:
5. Deep Insight: Why Does mBERT Win?
6. Conclusion & Future Outlook