Decoding the Cultural Spectator: Multi-Stage Profiling of Art Event Attendees via Social Media

Analysis of Online User Behaviour for Art and Culture Events

2017-01-01
Behnam Rahdari, Tahereh Arabghalizi, Marco Brambilla
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a domain-specific KDD pipeline for user behavioral profiling and interest prediction during large-scale art and culture events. Tested on "The Floating Piers" installation, the method utilizes LDA topic modeling and hierarchical clustering to categorize participants and decision trees to predict future engagement based on social media metadata.

TL;DR

This study presents a specialized data mining pipeline to profile and predict the interests of social media users attending major art events. By processing Twitter data from the massive "The Floating Piers" installation, the researchers successfully clustered over 20,000 users into three personas—Travel, Art, and Tech lovers—and developed a rule-based model to predict newcomer interests with 62% accuracy using minimal metadata.

Context & Motivation: Much More Than Just Keywords

In the era of "participatory art," events like Christo’s The Floating Piers generate millions of social interactions. However, for event organizers, a simple word count of "cool" or "beautiful" is insufficient. The challenge lies in understanding who the participants are and why they are there.

Existing research often hits two walls:

  1. Language Barriers: Most models focus exclusively on English, ignoring local contexts (e.g., Italian for an event in Lake Iseo).
  2. Superficial Similarity: Clustering by common words often groups people by event name rather than personal motivation.

The authors' insight was to treat the user as the document, aggregating biographies, hashtags, and tweets to reveal latent identities rather than transient reactions.

Methodology: The Behavioral Profiling Pipeline

The authors propose a structured Knowledge Discovery in Databases (KDD) process specifically tuned for social media's noisy environment.

1. Enrichment & Homogenization

Because Twitter doesn't provide gender or unified language, the authors used the NamSor API for gender inference and Yandex for translating all localized content into English. This ensures that a tweet in Italian and a tweet in French are mapped to the same semantic space.

2. High-Level Abstraction (LDA + PCA)

Instead of raw text, the model uses Latent Dirichlet Allocation (LDA) to calculate the probability of a user belonging to specific "topics." To handle correlation between topics and reduce noise, Principal Component Analysis (PCA) was applied, retaining only the top 3 components which captured 95% of the variance.

3. Clustering the Crowd

The study compared K-means, DBSCAN, and Hierarchical clustering. Experimental Results Comparison Table: Validation indices showing Hierarchical clustering on User Bios outperformed other configurations (higher Silhouette and lower Entropy).

Results: The Three Personas

The analysis identified three core segments through hierarchical dendrograms:

  • Travel Lovers (60%): Mostly locals (Italian) interested in the experience, food, and sightseeing.
  • Art Lovers (35%): A more international crowd, specifically visiting for the artistic value and photography.
  • Tech Lovers (5%): Social media managers and entrepreneurs interested in the logistics and digital marketing aspect of the event.

User Dendrogram Figure: The hierarchical dendrogram used to segment users based on their topic probabilities.

Predictive Insights: Profiling Newcomers

The final piece of the puzzle is the Decision Tree. The authors generated simple, interpretable rules to classify new users. Interestingly, the model found that the Biography Score (the latent topics in a user's profile) was the strongest predictor of their future engagement type.

  • Rule Example: If a user's Bio has a high "Art" topic score and they tweet in a language other than Italian, they are almost certainly an Art Lover.

Critical Perspective: The Power of Context

One of the paper's most salient points is that Biographies are superior to Tweets for profiling. While tweets tell you what the person is doing now, the bio tells you who they are.

Limitations:

  • Visual Gap: While the authors tracked Instagram locations, they did not perform image analysis. In contemporary art, the visual "vibe" often precedes the text.
  • Temporal Decay: Engagement peaked on day one and dropped sharply. The model doesn't fully account for how motivations might shift from early adopters to the general public.

Final Takeaway

For tourism boards and cultural institutions, this method offers a blueprint for predicting high-value visitors. By analyzing the "who" (Bio) rather than just the "what" (Tweet), organizations can tailor their communication strategies—marketing the process of art to the Tech Lover, and the beauty of the venue to the Art Lover.

Find Similar Papers

Try Our Examples

  • Search for recent studies on cross-lingual user profiling and sentiment analysis for large-scale public art exhibitions.
  • Which paper originally established the use of Latent Dirichlet Allocation (LDA) for user interest modeling in social networks, and how does this study's transition to a user-as-document approach differ?
  • Examine how the integration of multi-modal data, such as Instagram image processing, has improved the accuracy of user profiling compared to text-only pipelines like the one proposed here.
Contents
Decoding the Cultural Spectator: Multi-Stage Profiling of Art Event Attendees via Social Media
1. TL;DR
2. Context & Motivation: Much More Than Just Keywords
3. Methodology: The Behavioral Profiling Pipeline
3.1. 1. Enrichment & Homogenization
3.2. 2. High-Level Abstraction (LDA + PCA)
3.3. 3. Clustering the Crowd
4. Results: The Three Personas
5. Predictive Insights: Profiling Newcomers
6. Critical Perspective: The Power of Context
7. Final Takeaway