Beyond English: Building Bi-lingual Datasets for Roman-Urdu Social Media Analysis
Creation of Bi-lingual Social Network Dataset Using Classifiers
2014-01-01
Summary
Problem
Method
Results
Takeaways
Abstract
This paper introduces a robust framework for creating a topic-focused, bi-lingual social media dataset (English and Roman-Urdu) from Twitter. By employing a two-stage classification process involving keyword matching and NLP-based lexical analysis, the authors curated a 82,224-tweet corpus, achieving a 96% language classification accuracy.
## TL;DR
Social media is a multilingual hub, yet many analytical studies focus solely on English. This paper presents a methodology for creating a high-quality, topic-focused dataset targeting **Roman-Urdu** (Urdu written in English script) and English. By implementing a dual-stage classification system, the researchers achieved **96% accuracy** in distinguishing languages, capturing over 21,000 tweets that would otherwise have been discarded as "noise."
## Background: The Hidden Value of Localized Scripts
Twitter has become a digital pulse for public behavior, but in countries like Pakistan, English is not the sole medium of expression. Approximately 8 million active users communicate in Roman-Urdu. This phonetic script lacks standardized spelling, making it invisible to traditional NLP tools. The authors argue that to accurately predict political outcomes or consumer sentiment in these regions, we must first build robust, classified datasets that recognize these local variations.
## Methodology: The Two-Fold Classification Engine
The researchers didn't just collect data; they built a pipeline to refine it through two critical iterations.
### 1. Subject Classification (Noise Removal)
Twitter data is notoriously messy. Political hashtags are often hijacked by real estate spammers (e.g., "Bahria Town", "Kanal", "DHA"). The first stage used a keyword-based subject classifier to purge non-political tweets, ensuring the resulting corpus was strictly about the political landscape.
### 2. Bi-lingual Language Classification
How do you tell English and Roman-Urdu apart when they use the same alphabet? The authors proposed a weight-based mechanism:
* **Tokenization**: Stripping digits, hyperlinks, and special characters.
* **Lexical Filtering**: Comparing tokens against a **2,000-word English lexicon** and **WordNet**.
* **Weight Calculation**: A tweet is assigned a weight based on the ratio of English tokens ($W_{eng}$) to total tokens ($W_{Total}$):
$$Weight_{tweet} = \frac{W_{eng}}{W_{Total}} * 100$$
The team empirically determined that a **35% threshold** provided the optimal balance for classification.

*Figure 1: The proposed classification system workflow.*
## Experimental Results & Performance
The system was tested against 89,000 tweets collected over four months.
* **Subject Classifier**: Acheived 98% precision in topic relevance, removing roughly 6,800 noisy tweets.
* **Language Classifier**: The 35% threshold yielded the best metrics across the board:
* **Accuracy**: 96%
* **F-Score**: 97%
* **Sensitivity**: 99%

*Table 1: Evaluation of the classifier at different weight thresholds.*
The final dataset breakdown showed that **26.4%** of the political discourse was in Roman-Urdu. Without this bilingual approach, a quarter of the public's voice would have been ignored.
## Critical Insight: Why This Matters
The brilliance of this work lies in its simplicity. Instead of training complex, data-hungry deep learning models (which were less accessible at the time of publication), the authors used a **lexicon-driven weighting system** that accounts for the inherent "semantic fluctuations" of Romanized scripts.
### Limitations and Future Outlook
While highly effective, the model faces challenges with "Code-Switching"—tweets where users mix English and Roman-Urdu in a single sentence. The authors note that the next step is to perform **intra-tweet classification** to handle these hybrid expressions more precisely.
As global social media analysis moves toward total inclusion, this work serves as a foundational blueprint for handling low-resource, informally written languages across the digital landscape.
