Bridging the Gap: Mapping Social Media "Short-Names" to Real-World Locations

Identifying Short-Names for Place Entities from Social Networks

2017-10-26
Faizan Wajid, Hong Wei, Hanan Samet
Summary
Problem
Method
Results
Takeaways
Abstract

Identifying Short-Names for Place Entities is a study focused on mapping colloquial abbreviations and nicknames (short-names) from social networks to official organizations. The authors propose a heuristic-based generation framework and a TF-IDF scoring mechanism to resolve ambiguous spatial queries into precise geographic coordinates.

TL;DR

In the world of social media, official names like the "University of Maryland" often take a backseat to colloquialisms like "UMD" or "UMD CS." This paper presents an automated framework to identify these short-names from social networks, generate potential variants using five heuristic strategies, and link them to precise geographic coordinates. It is a critical step toward making spatial search engines as fluent in slang as their users.

The "Toponym Resolution" Problem

Most geographic databases (like GeoNames) are static entities. They store official names and perhaps a few recognized aliases. However, human language—especially on platforms like Twitter (X)—is fluid. A user searching for "GDubs" likely wants George Washington University, but if a search engine hasn't mapped that specific slang, the query fails.

The core challenge is two-fold:

  1. Variation: Organizations are referred to by abbreviations, syllable-based nicknames, or even rearranged titles.
  2. Ambiguity: A short-name like "UM" could refer to many different universities depending on context and proximity.

Methodology: The Five Heuristic Buckets

The authors categorize short-name generation into five logical "buckets" (B1–B5). This systematic approach allows the system to predict how a name might be shortened before even seeing it in the wild.

  • B1: Initializations: Standard acronyms (e.g., U.S.A.).
  • B2: State Abbreviations: Incorporating "MD" for Maryland or "VA" for Virginia.
  • B3: Word Swapping: "State Department" vs. "Department of State."
  • B4: Common Abbreviations: Using "Univ" for University or "C2" for Community College.
  • B5: Syllables: Taking the first syllable, such as "U.Mich" for Michigan.

Model Architecture Placeholder Figure 1: Conceptual overview of identifying entities from social streams.

To validate these, the researchers scraped 3,200 tweets from selected organizations in the D.C.-Maryland-Virginia (DMV) area. They used TF-IDF (Term Frequency-Inverse Document Frequency) to score the relationship between a short-name and an organization, ensuring that the most "renowned" entity gets priority in local searches.

Experimental Insights

The study analyzed 30 organizations across four sectors: Education, Commercial, Government, and Military.

Bucket Performance Comparison Table 1: Distribution of short-name types across different organization categories.

Key Findings:

  • Education entities are the most "creative," with high usage of both initializations (68.3%) and common abbreviations (70.3%).
  • Commercial entities rely almost exclusively on initializations (88.8%).
  • Emergent Slang: The system successfully flagged unmapped terms like "TheFeddy" (Federal Reserve), showcasing the power of social media monitoring over manual database entry.

Critical Analysis & Future Outlook

While the heuristic approach is robust, it has limitations. The authors acknowledge that as the scope moves from a local area (like the DMV) to a global scale, the "collision" of short-names (e.g., "UM" could be Michigan, Maryland, or Miami) becomes a major hurdle.

The Takeaway: This research highlights that spatial querying must move beyond the "official name." By treating social media as a dynamic training set, developers can build GIS systems that feel more intuitive. Future work involving Sentiment Analysis and broader platform integration (Facebook, Yelp) will likely refine these mappings, making "GDubs" as searchable as "George Washington University."

Short-Name Database Demo An example of the final output: mapping a short-name to its official entity and coordinates.

Find Similar Papers

Try Our Examples

  • Find recent research papers that utilize Twitter or social media data for dynamic Gazeteer enhancement or automated Toponym expansion.
  • Which studies first explored the use of TF-IDF or similar weighting schemes specifically for disambiguating short-names in spatial entity recognition?
  • Explore the application of Large Language Models (LLMs) in replacing heuristic-based short-name generation for geographic information systems.
Contents
Bridging the Gap: Mapping Social Media "Short-Names" to Real-World Locations
1. TL;DR
2. The "Toponym Resolution" Problem
3. Methodology: The Five Heuristic Buckets
4. Experimental Insights
5. Critical Analysis & Future Outlook