Beyond English: Building Bi-lingual Datasets for Roman-Urdu Social Media Analysis

Creation of Bi-lingual Social Network Dataset Using Classifiers

2014-01-01
Iqra Javed, Hammad Afzal
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a robust framework for creating a topic-focused, bi-lingual social media dataset (English and Roman-Urdu) from Twitter. By employing a two-stage classification process involving keyword matching and NLP-based lexical analysis, the authors curated a 82,224-tweet corpus, achieving a 96% language classification accuracy.

    ## TL;DR
    Social media is a multilingual hub, yet many analytical studies focus solely on English. This paper presents a methodology for creating a high-quality, topic-focused dataset targeting **Roman-Urdu** (Urdu written in English script) and English. By implementing a dual-stage classification system, the researchers achieved **96% accuracy** in distinguishing languages, capturing over 21,000 tweets that would otherwise have been discarded as "noise."

    ## Background: The Hidden Value of Localized Scripts
    Twitter has become a digital pulse for public behavior, but in countries like Pakistan, English is not the sole medium of expression. Approximately 8 million active users communicate in Roman-Urdu. This phonetic script lacks standardized spelling, making it invisible to traditional NLP tools. The authors argue that to accurately predict political outcomes or consumer sentiment in these regions, we must first build robust, classified datasets that recognize these local variations.

    ## Methodology: The Two-Fold Classification Engine
    The researchers didn't just collect data; they built a pipeline to refine it through two critical iterations.

    ### 1. Subject Classification (Noise Removal)
    Twitter data is notoriously messy. Political hashtags are often hijacked by real estate spammers (e.g., "Bahria Town", "Kanal", "DHA"). The first stage used a keyword-based subject classifier to purge non-political tweets, ensuring the resulting corpus was strictly about the political landscape.

    ### 2. Bi-lingual Language Classification
    How do you tell English and Roman-Urdu apart when they use the same alphabet? The authors proposed a weight-based mechanism:
    *   **Tokenization**: Stripping digits, hyperlinks, and special characters.
    *   **Lexical Filtering**: Comparing tokens against a **2,000-word English lexicon** and **WordNet**.
    *   **Weight Calculation**: A tweet is assigned a weight based on the ratio of English tokens ($W_{eng}$) to total tokens ($W_{Total}$):
    
    $$Weight_{tweet} = \frac{W_{eng}}{W_{Total}} * 100$$
    
    The team empirically determined that a **35% threshold** provided the optimal balance for classification.

    ![System Workflow](https://cdn.atominnolab.com/wisdoc/images/20260523-08d308c5-ac54-4700-a75f-5825031869cc/page_005_block_002.png)
    *Figure 1: The proposed classification system workflow.*

    ## Experimental Results & Performance
    The system was tested against 89,000 tweets collected over four months. 

    *   **Subject Classifier**: Acheived 98% precision in topic relevance, removing roughly 6,800 noisy tweets.
    *   **Language Classifier**: The 35% threshold yielded the best metrics across the board:
        *   **Accuracy**: 96%
        *   **F-Score**: 97%
        *   **Sensitivity**: 99%

    ![Threshold Comparison](https://cdn.atominnolab.com/wisdoc/tables/20260523-08d308c5-ac54-4700-a75f-5825031869cc/page_006_block_008.png)
    *Table 1: Evaluation of the classifier at different weight thresholds.*

    The final dataset breakdown showed that **26.4%** of the political discourse was in Roman-Urdu. Without this bilingual approach, a quarter of the public's voice would have been ignored.

    ## Critical Insight: Why This Matters
    The brilliance of this work lies in its simplicity. Instead of training complex, data-hungry deep learning models (which were less accessible at the time of publication), the authors used a **lexicon-driven weighting system** that accounts for the inherent "semantic fluctuations" of Romanized scripts. 

    ### Limitations and Future Outlook
    While highly effective, the model faces challenges with "Code-Switching"—tweets where users mix English and Roman-Urdu in a single sentence. The authors note that the next step is to perform **intra-tweet classification** to handle these hybrid expressions more precisely.

    As global social media analysis moves toward total inclusion, this work serves as a foundational blueprint for handling low-resource, informally written languages across the digital landscape.

Find Similar Papers

Try Our Examples

  • Find recent papers or SOTA methods specifically designed for Romanized Urdu (Roman-Urdu) sentiment analysis and text classification to see how methodology has evolved since 2013.
  • Which paper first formally defined the linguistic challenges of "Roman-Urdu" in social media, and how does this paper's token-weighting method compare to early character n-gram models?
  • Identify studies that have applied similar bilingual lexical weighting techniques to other Romanized languages, such as Roman-Hindi or Hinglish, in the context of political opinion mining.
Contents
Beyond English: Building Bi-lingual Datasets for Roman-Urdu Social Media Analysis
1. TL;DR
2. Background: The Hidden Value of Localized Scripts
3. Methodology: The Two-Fold Classification Engine
3.1. 1. Subject Classification (Noise Removal)
3.2. 2. Bi-lingual Language Classification
4. Experimental Results & Performance
5. Critical Insight: Why This Matters
5.1. Limitations and Future Outlook