Re-evaluating Event Discovery: A High-Specificity Approach to Facebook Recommendations
Recommendation system for Facebook public events based on probabilistic classification and re-ranking
The paper introduces a two-stage recommendation framework for Facebook public events, combining a probabilistic ensemble classifier (Random Forest + Logistic Regression) with a re-ranking layer. It focuses on solving the intense cold-start problem of short-term events by prioritizing the elimination of unsuitable items (negative filtering) over traditional similarity-based ranking.
TL;DR
Unlike movies or books, Facebook Events are ephemeral, making traditional Collaborative Filtering obsolete. This paper presents a two-stage system that treats recommendation as a Probabilistic Classification problem followed by Diversity-Aware Re-ranking. By optimizing for the removal of "unsuitable" events—rather than just finding the most similar ones—the system achieves higher reliability and better user experience in cold-start scenarios.
Problem & Motivation: The "Short-Lived" Dilemma
In a world dominated by static content (Netflix, Amazon), Events represent a unique challenge. They exist for a fleeting window of time, often have no prior rating history, and are subject to strict geographical constraints.
The authors argue that Facebook's environment is distinct from "Event-Based Social Networks" (EBSNs) like Meetup. On Meetup, group membership is public and driving interest; on Facebook, users are often strangers, and privacy settings obscure social ties. Therefore, the system must rely on Content, Location, Time, and Ownership to bridge the gap.
Methodology: The Two-Stage Framework
1. Feature Engineering: The "Big Four"
The model constructs a rich feature vector for every user-event pair :
- Content: TF-IDF vectors processed with
vnTokenizerfor Vietnamese support. - Location: Latitude/Longitude distances from the user's historical attendance.
- Time: Ratios of event occurrences on specific days of the week.
- Owner: Track record and reputation of the event organizer.
2. Stage I: High-Specificity Classification
The core innovation lies in the Optimization Procedure. Instead of a simple 0.5 threshold, the authors use a support vector machine (SVM) to find an optimal threshold for the negative probability . This ensures that if the system isn't sure, it errs on the side of not showing a potentially annoying event.
Insight: Notice how Location (max/avg dist) is the dominant factor—users are highly sensitive to travel distance.
3. Stage II: Re-ranking for Diversity
Once the "bad" events are filtered out, the remaining candidates are re-ordered. The authors compared five techniques:
- Similarity
- Reverse Similarity (to boost discovery)
- Serendipity
- Popularity
- Relative Popularity (Final choice)
Experiments & Results
The study was conducted on a dataset of 1,160 public events in Ho Chi Minh City.
The Trade-off: Accuracy vs. User Trust
By prioritizing Specificity (Scenario 3), the authors achieved a 12% lead over other probabilistic models. While Accuracy and Recall are standard, in recommendation, "User Trust" is built by reducing false positives—events the user explicitly dislikes.

The Relative Popularity scoring technique outperformed others in MAP@10 (Mean Average Precision). This technique uses the ratio of positive reactions to total interactions, acting as a biological filter for event "quality" or "vibe."
Critical Analysis & Conclusion
Takeaway
The paper successfully demonstrates that for Facebook Public Events, location is the primary filter, followed by content. The two-stage architecture provides a safety net: the classifier protects against irrelevant content, while the re-ranker ensures the top 10 results aren't just a repetitive wall of identical items.
Limitations & Future Work
- Data Scale: The study is localized to Ho Chi Minh City. Scaling this to global data might reveal different cultural "Time" or "Owner" patterns.
- Implicit Feedback: The current model relies on explicit tags (Attending, Interested, Declined). Incorporating implicit signals (dwell time, clicks) would likely improve the classification nuance.
This work serves as a foundational blueprint for developers building event-based apps on social platforms where social graphs are restricted.
