Stanford-TUN v2.0: Bridging the Gap in Tunisian Arabic Social Media Parsing

Treebank Creation and Parser Generation for Tunisian Social Media Text

2020-11-01
Asma Mekki, Inès Zribi, Mariem Ellouze, Lamia Hadrich Belguith
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces the creation of a specialized Treebank and a semi-automatic parser for Tunisian Arabic (TA) specifically tailored for Social Media Text. Leveraging a modified Stanford Parser (Stanford-TUN v2.0), the authors address the linguistic idiosyncrasies of user-generated content, achieving a state-of-the-art (SOTA) F-measure of 83.32% for this low-resource dialect.

TL;DR

Researchers from the University of Sfax have developed a specialized Treebank and Parser for Tunisian Arabic (TA) social media text. By bootstrapping an existing parser and refining it with a hand-annotated corpus of 10,000 sentences, they achieved an 83.32% F-measure, significantly outperforming baseline models trained on intellectualized or standard Arabic dialects.

Background & Positioning

In the landscape of Arabic NLP, Modern Standard Arabic (MSA) has historically received the lion's share of attention. However, everyday digital communication—accelerated by the COVID-19 pandemic—happens in Tunisian Arabic (TA). TA is morphologically rich, syntactically flexible, and, until now, lacks the robust treebanks required for deep syntactic analysis. This work moves TA from a "low-resource" status toward a more mature NLP ecosystem by tackling the specific chaos of social media text.

The Social Media Problem: Why MSA Parsers Fail

Social media dialect (SMD) isn't just "informal" Arabic; it’s a distinct linguistic beast characterized by:

  • Orthographic Chaos: A single word like twa (now) can be written in nine different ways.
  • Non-Grammatical Tokens: Emojis, onomatopoeia (e.g., "hhh" for laughter), and foreign inclusions (French/Arabizi).
  • Complex Negation: TA uses unique circumfix negation (e.g., ma...sh) that MSA-based models struggle to segment and link correctly.

Methodology: Bootstrapping and Tagset Extension

The authors didn't start from scratch. They utilized a bootstrapping method to expedite the labor-intensive process of treebank creation:

  1. Preprocessing: Manual tokenization and normalization using the CODA-TUN convention.
  2. Lexicon Integration: Adding 38 onomatopoeias and 51 emojis directly into the parser's knowledge base.
  3. Tagset Expansion: Introducing specific tags like ONM (Onomatopoeia), EMJ (Emoji), and RPQ (Question Particle) under a new parent tag NG (Non-Grammatical).

Overall Architecture/Negation Comparison Fig 1: A comparison between a generic model (Left) and the new SMD model (Right), showing superior handling of TA negation structures.

Experimental Insights: Specificity Beats Scale

The team conducted five experiments comparing different training sets (SMD, Intellectualized Dialect, and Spontaneous Dialect).

Key Findings:

  • Mixing isn't always better: Using a model trained only on SMD performed better for social media text than a model trained on all dialect forms. The "ALL" model often misclassified proper nouns (NNP) as common nouns (NN).
  • Analysis Methods: While PCFG (Probabilistic Context-Free Grammar) is a standard, Dependency Analysis proved more robust for the longer, punctuation-poor sentences typical of Facebook comments.

Experimental Results Comparison Table 1: Performance metrics showing the SMD-specific model reaching ~80% F-measure compared to the ~53% baseline.

Deep Insight & Conclusion

The success of this work lies in its linguistic sensitivity. Rather than treating social media noise as something to be "cleaned away," the authors embraced it as a syntactic component. By assigning formal tags to emojis and onomatopoeias, the parser maintains the structural integrity of the sentence, allowing for better downstream tasks like Sentiment Analysis.

Limitations: The reliance on manual tokenization remains a bottleneck. Future Work: The development of an automatic tokenization system for TA and a more formalized TA-specific grammar are the next frontiers for the MIRACL Laboratory team.

Sample Parsed Sentences Fig 2: Golden trees showing correct handling of imperative verbs and coordination in Tunisian social media posts.

Find Similar Papers

Try Our Examples

  • Search for recent studies applying Universal Dependencies (UD) specifically to Arabic dialects in social media contexts.
  • Which paper first proposed the CODA-TUN orthographic convention, and how has it evolved to accommodate Arabizi and French code-switching?
  • Explore how the parenthetic tag "NG" (non-grammatical) introduced in this paper is handled in other low-resource dialectal NLP tasks like Sentiment Analysis or Named Entity Recognition.
Contents
Stanford-TUN v2.0: Bridging the Gap in Tunisian Arabic Social Media Parsing
1. TL;DR
2. Background & Positioning
3. The Social Media Problem: Why MSA Parsers Fail
4. Methodology: Bootstrapping and Tagset Extension
5. Experimental Insights: Specificity Beats Scale
6. Deep Insight & Conclusion