Venue Semantics: Bridging the Modality Gap in Social Media Topic Modeling
Venue Semantics: Multimedia Topic Modeling of Social Media Contents
This paper introduces a Multimedia Venue Topic Modeling (MVTM) framework designed to extract semantic venue topics from heterogeneous and independent social media data. By leveraging cross-platform sources like Flickr, Wikitravel, and TripAdvisor, the method bridges the gap between disconnected Foursquare images and texts, achieving SOTA performance in semantic-based venue summarization.
TL;DR
Understanding the "vibe" of a physical venue (like Marina Bay Sands) through social media is difficult because users upload photos and short tips independently. This paper proposes a Multimedia Venue Topic Modeling (MVTM) framework that uses cross-platform data (Flickr, Wikitravel) as a "semantic bridge" to connect independent texts and images. By treating topic discovery as a constrained dense subgraph detection problem, it generates rich, multimedia summaries that far exceed the quality of traditional text-only methods.
The Challenge: The Independence of Modalities
In standard LBSNs like Foursquare, a user might post a photo of their hotel room without a caption, while another user leaves a text tip about the rooftop pool without a photo.
- Heterogeneity: Data types range from short, noisy "tips" to unlabelled images.
- Independence: Unlike Flickr (where tags describe photos), Foursquare content lacks explicit links between a specific image and a specific text.
- Noise: Social media is full of irrelevant "portraits" (selfies) and spam tips that dilute the venue's actual semantic identity.
Methodology: Cross-Platform Enrichment & Constrained Graph Shift
The authors' core insight is that while Foursquare data is disconnected, other platforms can provide the missing links.
1. Data Enrichment & Noise Reduction
- The Flickr Bridge: Since Flickr photos have both visual content and tags, they act as an intermediary. Foursquare images are linked to Flickr images via visual similarity, and Foursquare texts are linked to Flickr tags via semantic similarity.
- Textual/Visual Cleaning: They use LDA to filter textual noise and a Part-based detection model (typically used in CV for object detection) to remove "noisy" photos—specifically selfies where faces or bodies occupy >45% of the frame.
2. Multi-modality Graph Construction
The system builds a graph where nodes include latent textual topics (), Foursquare photos (), and Flickr photos ().
- Visual Similarity: Calculated using Local Binary Patterns (LBP) and cosine similarity.
- Textual Similarity: Based on the posterior probability of words in an LDA model.
Figure: The three-stage framework: Data Enrichment, Graph Construction, and Topic Modeling.
3. Dense Subgraph Detection (Graph Shift)
To find a "topic," the authors search for the densest subgraphs. However, they add a critical constraint: each topic must contain at least one textual node. They solve this by optimizing a Lagrangian function under the constraint , ensuring the resulting multimedia topic is anchored by a semantic concept.
Experimental Validation
The model was tested on 41 popular venues in Singapore. Instead of just seeing a list of words, users see "Semantic Summaries." For example, at the Esplanade, the model identifies distinct topics like "Architecture/Landmarks" and "Cultural Activities (Books/Arts)," each populated with the most relevant images and keywords.
Figure: Examples of semantic summarization for venues like Marina Barrage and The Esplanade.
Quantitative Performance
- Satisfaction: Multimedia summaries scored 8.5/10, compared to 7.4/10 for text-only versions.
- Consistency: The graph-based approach successfully merged different visual angles of the same landmark into a single semantic topic, a feat traditional visual-only clustering fails to achieve.
Critical Insight & Future Outlook
This paper proves that the "semantic meaning" of a place isn't found in one modality alone. The real value is in the relational structure between modalities. While modern LLMs and CLIP-based models (which didn't exist when this was written) now handle multimodal embeddings more elegantly, this paper’s focus on graph-based dense subgraph detection remains a robust way to handle the noise and independence inherent in social media.
Future Directions: The authors suggest incorporating "User Influence"—recognizing that a tip from a local expert should carry more weight than a one-time tourist's selfie in defining a venue's semantics.
