Probabilistic Lexicon Expansion: Bridging the Gap in Social Media Slang

A probalistic approach to automatically extract new words from social media

2016-08-18
Geetika Sarna, M. P. S. Bhatia
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a probabilistic framework for the automatic extraction of new keywords from short-form social media content (specifically Twitter). By leveraging Bayes Rule and Jaro-Winkler distance, the method identifies links between unknown terms and existing thematic dictionaries to dynamically update lexicons for community detection.

TL;DR

Social media moves faster than any dictionary. To solve this, the authors propose a probabilistic framework that automatically identifies "new" words (neologisms) from short-form messages like Tweets and maps them to specific communities (e.g., Education, Bullying, Terrorism). By using conditional probability and graph-based link pruning, they turn static lexicons into dynamic graphs that evolve alongside user behavior.

Context & Motivation

In the world of Natural Language Processing (NLP), micro-blogs are a nightmare. Standard tools built for research papers or news articles fail because:

  1. Length Constraints: A 140-character tweet doesn't have enough "surface area" for traditional TF-IDF or positional features.
  2. Vocabulary Volatility: Slang, abbreviations, and coded language (especially in malicious communities) appear and change almost daily.

The authors argue that identifying new keywords is not just about detection; it's about association. If a new word frequently appears alongside known "Bullying" terms, the system should learn to treat that new word as part of the bullying lexicon.

Methodology: The Bayes-Jaro Pipeline

The proposed framework follows a sophisticated four-stage process:

1. Linguistic Pre-processing

Using the Stanford POS Tagger, the system strips away the "noise" (prepositions, adverbs) and focuses on the "signal" (nouns, adjectives, verbs). This is followed by stemming to reduce words to their base forms (e.g., "fishing" to "fish").

2. Similarity Matching with Jaro-Winkler

How do you know if a word is truly "new"? The system compares tokens against existing dictionaries using Jaro-Winkler Distance. This is particularly effective for short strings because it gives more weight to the prefix of the string, making it resilient to slight spelling variations common on Twitter.

3. Probabilistic Graph Construction

This is the core "intelligence" of the paper. Instead of just counting frequencies, the authors build a directed graph where:

  • Nodes: Both existing () and new keywords ().
  • Edges: Determined by conditional probability —the likelihood of a new word appearing given a known keyword.

Proposed Framework

4. Chi-Square Link Pruning

Not every pairing is meaningful. To separate true associations from random noise, the authors apply a Chi-Square test. If the dependency between two words is statistically significant (p < 0.05), the link is retained; otherwise, it is pruned.

Experiments & Real-World Results

The researchers tested their approach on 1,000 live tweets, generating a pool of roughly 2,900 unique keywords.

Key Benchmarks:

  • Jaro-Winkler Accuracy: 92.5%. The system is excellent at distinguishing known terms from unknown ones.
  • Domain Assignment Accuracy: 74.4%. While solid, this reflects the difficulty of the task. Errors often occurred when a word was used across multiple domains, making it a "general" word rather than a domain-specific one.

Graph After Pruning In the figure above, the word "fat" was statistically linked to "hate" and "mad," allowing the "Bullying" dictionary to be automatically updated with this "new" context.

Critical Insight & Future Outlook

The beauty of this approach is its Domain Independence. Whether you are tracking sports trends or monitoring online radicalization, the underlying math remains the same.

Limitations to Consider:

  • Data Sparsity: The 74.4% accuracy suggests that 1,000 tweets are not enough. Probabilistic models thrive on Big Data; a larger corpus would likely reduce the error rate.
  • Context Shifting: A word might mean something in the "Sports" community today and something entirely different in a "Political" community tomorrow.

The Takeaway: This research provides a vital blueprint for building living lexicons. In the future, combining this probabilistic approach with deep learning (like Transformer-based embeddings) could create even more robust systems for real-time social media intelligence.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize Graph Neural Networks (GNNs) for keyword extraction and expansion in short-text social media data.
  • Which original paper established the Jaro-Winkler distance metric, and how have modern adaptions improved its accuracy for entity resolution in informal text?
  • How are large language models (LLMs) currently being used to perform zero-shot or few-shot expansion of thematic lexicons in social network analysis?
Contents
Probabilistic Lexicon Expansion: Bridging the Gap in Social Media Slang
1. TL;DR
2. Context & Motivation
3. Methodology: The Bayes-Jaro Pipeline
3.1. 1. Linguistic Pre-processing
3.2. 2. Similarity Matching with Jaro-Winkler
3.3. 3. Probabilistic Graph Construction
3.4. 4. Chi-Square Link Pruning
4. Experiments & Real-World Results
5. Critical Insight & Future Outlook