Your Friends Reveal You're Online: Predicting OSN Availability through Social Ties
Evaluating the Impact of Friends in Predicting User’s Availability in Online Social Networks
The paper investigates predicting a user's online availability (online/offline status) in Online Social Networks (OSNs) specifically by leveraging the real-time status of their friends. Using a custom Facebook dataset and various machine learning classifiers, the study achieves an AUC of 0.803, demonstrating that social ties provide significant predictive signals for individual behavior.
TL;DR
Is your privacy safe just because you hide your "Last Seen" status? Not necessarily. This paper proves that the online status of your friends is a strong predictor of your own availability. By analyzing a real-world Facebook dataset of over 60,000 users, researchers demonstrated that machine learning models—specifically Decision Trees and Random Forests—can infer a user's presence with an AUC of over 0.80 based solely on their friends' activity.
Academic Context: This work shifts the focus from purely temporal individual history to social-contextual prediction, highlighting a significant "side-channel" for privacy leakage in Online Social Networks (OSNs).
The "Invisible" Problem: Social Correlation
Most existing research on user availability (uptime) treats users as isolated islands, predicting future status based on past patterns (e.g., "User A is usually online at 9 PM"). However, humans are social animals. We join OSNs to interact.
The authors identify a critical gap: Temporal Homophily. If your close friends are online, you are much more likely to be online as well. This creates a privacy risk—even if YOU hide your status, your friends might inadvertently leak it.
Methodology: Turning Friends into Features
The researchers developed a dedicated Facebook application to monitor the chat status of 204 registered users and their cumulative 66,880 friends.
1. The Dataset
- Duration: 32 consecutive days.
- Granularity: 5-minute sampling intervals (288 time steps per day).
- Feature Set: A simple yet effective triplet:
<onFriends, offFriends, S>, whereSis the target status (0 for offline, 1 for online).
2. Model Selection
The study benchmarked several classical and ensemble learning algorithms to see which could best handle the non-linear relationship between social groups and individual behavior:
- C4.5 Decision Tree: Built a hierarchical logic based on friend counts.
- Random Decision Forests: Combined multiple trees to reduce variance.
- Probabilistic Neural Network (PNN): Used Gaussian functions for dynamic classification.
- k-Nearest Neighbor (k-NN): Looked at historically similar friend-status distributions.
Table: The core attributes used for prediction.
Experimental Insights
The data revealed a clear cyclic day/night pattern, with peaks during lunch and evening. Interestingly, the researchers found that most users are offline the majority of the time, creating a "Class Imbalance" problem.
Performance Highlights:
- Top Performers: C4.5 Decision Trees achieved the highest accuracy (63.7%), while Random Forests and k-NN tied for the best AUC (0.803).
- The Precision Win: All models showed high Precision (~0.82), meaning when the model predicts a user is "Online," it is usually correct.
- The Sensitivity Challenge: Sensitivity was lower (~0.35), largely because the "Offline" class is so dominant that models struggle to catch every single brief "Online" session.
Table: Comparison of Accuracy, Kappa, and AUC.
Deep Insight: Why Why Does This Matter?
The value of this research lies in its implications for Privacy and System Design:
- Privacy Vulnerability: An adversary doesn't need to be your friend to know when you're online. They only need to see the public status of a few of your friends to build a probabilistic model of your habits.
- Notification Efficiency: OSN providers can use these models to determine the optimal time to send push notifications, ensuring they land exactly when a user is likely to be responsive.
- Data Caching: In decentralized networks, knowing when a "cluster" of friends will be online allows for smarter data replication and energy saving.
Conclusion & Future Outlook
The study successfully validates that availability is a collective behavior, not just an individual one.
Limitations: The model is currently "identity-blind"—it treats all friends as equal. In reality, a partner's status likely has more predictive weight than a distant high school acquaintance. Next Steps: Future research should incorporate edge weights (how close is the friend?) and Sequence Modeling (using LSTMs or Transformers) to capture how availability flows through a social graph over time.
Takeaway for the Industry: "Invisible" modes are not enough. True privacy in social networks requires protecting not just the user, but the metadata of the user's social environment.
