From Geo-Coordinates to Sentiment: Enriching City POIs with Social Intelligence
On the enrichment of a RDF repository of city points of interest based on social data
This paper presents a framework to enrich RDF repositories of city Points of Interest (POIs) by integrating unstructured data from Social Networking Sites (SNSs) like Foursquare and Yelp. The core methodology combines a hybrid string-matching technique for entity reconciliation with a pattern-based NLP approach for fine-grained sentiment extraction and structured repo enrichment.
TL;DR
Researchers have developed a method to bridge the gap between "stale" structured geographic data and "vibrant" social media reviews. By combining a new similarity metric for cross-platform entity matching with NLP-based sentiment extraction, they can automatically enrich RDF repositories with qualitative insights—telling you not just where a restaurant is, but if the "service is kind" or the "seafood is fresh."
The Data Gap: Why "Location" Isn't Enough
In the era of Digital Cities, simply knowing the GPS coordinates of a museum doesn't suffice. Current RDF repositories, often harvested from sources like Google Fusion Tables, suffer from the "Sparsity Problem": they have the name and the street, but they lack the soul of the place.
The authors argue that Social Networking Sites (SNSs) like Yelp and Foursquare are the solution. However, mapping a database entry to a social page is a nightmare of "string noise"—one platform lists "The Louvre," another "Musée du Louvre." Standard exact-matching fails, and existing entity reconciliation methods are rarely optimized for the unique constraints of geographic POIs.
Hybrid Similarity: Character vs. Word Logic
The authors' first major contribution is a similarity formula that addresses the "Edit Distance" vs "Token Overlap" trade-off.
- Character-level (Levenshtein): Great for catching typos or minor suffix changes.
- Word-level (Jaccard): Excellent for phrases where word order changes but meaning remains (e.g., "Hotel Baltum" vs "Baltum Hotel").
By averaging these and adding a "Geographic Safety Net" (where names can be slightly less similar if the GPS coordinates are very close), they achieved a significant boost in precision.

Mining the "Why": Pattern-Based NLP
Unlike basic sentiment analysis that just gives a "thumbs up/down," this method extracts why people feel a certain way using five linguistic patterns:
- Pattern 1 & 2 (Object-Attribute): Identifies qualities like "Great food" or "The sandwich is good."
- Pattern 3, 4, & 5 (Feeling/Advice): Captures personal advice like "I suggest you visit the basement."
The result is a structured "Notation" (score) from -10 to 10 and three clear enrichment categories: General Assessment, Tips, and Specific Ideas.
Experimental Results
Testing the approach on landmarks like the Louvre Museum showed that the system could effectively filter through thousands of words to find specific complaints (e.g., "massive and uncomfortable") and specific praises (e.g., "panoramic view").

Critical Insight & Future Outlook
While the pattern-based approach is a strong starting point, the paper acknowledges a major limitation: Lexicon Dependency. Currently, if an adjective isn't in their pre-defined list of ~3,000 words, the system stays silent.
The future of this work lies in SentiWordNet integration and deeper linguistic nuance (handling adverbs like "not too bad" correctly). For city planners and app developers, this research maps out the pipeline for turning messy social chatter into high-value, machine-readable Knowledge Graphs.
Conclusion
The enrichment of POIs turns a map into a guide. By moving beyond coordinates and into the realm of structured social insights, we move closer to search engines that can answer complex queries like "Where is a seafood restaurant in Paris with kind service?"
