Inferring Social Ties: Do Your Twitter Follows Reveal Your Real Friends?
Inferring social ties from common activities in twier
This paper investigates the feasibility of inferring social ties between Twitter users by analyzing their shared activities, specifically the common celebrity profiles they follow. The authors evaluate the predictive power of two established probabilistic models—Kossinets & Watts and Crandall et al.—on a dataset of over 21,000 active users.
TL;DR
Can the list of celebrities you follow on Twitter reveal who your real-life friends are? This research explores the link between shared interests and social ties. By testing two classical probabilistic models on a dataset of 21,000 users, the authors found that following the same "niche" celebrities (like CEOs or Authors) is a surprisingly accurate predictor of friendship, whereas general celebrity following follows a much more chaotic, less predictable pattern.
Contextual Positioning
In the landscape of social network analysis, this paper serves as a critical bridge between behavioral coincidence and relationship inference. It validates that even "weak" signals—such as following the same public figures—carry enough structural information to de-anonymize or predict "strong" social ties, particularly when those signals are categorized by professional or intellectual interests.
The Core Challenge: Privacy in Public Data
Online behavior is often public, while ties (friendships) are frequently private. Previous work by Kossinets & Watts used email logs, and Crandall et al. used GPS data (Flickr) to infer ties. The challenge here is asymmetry: Twitter is a directed graph where following a celebrity doesn't imply a personal bond, yet the "coincidence" of two people following the same subset of celebrities might suggest they know each other or belong to the same offline community.
Methodology: Two Models, One Question
The authors leveraged a clever temporal split:
- Ground Truth: Friendships (defined as reciprocal follows) established by July 2009.
- Activities: Celebrity follows initiated after July 2009 (up to 2013).
This avoids the "circularity" problem where friends follow a celebrity because their friend recommended it, focusing instead on whether shared interest predicts pre-existing social bonds.
The Two Probability Lenses:
- Kossinets & Watts Model: Assumes activities are independent and provides a "baseline" or upper bound. If the actual friendship probability stays below this line, the model is considered a valid bound.
- Crandall et al. Model: A more ambitious model that attempts to provide a "realistic" estimate based on a hypergeometric-style distribution of common items.
Figure 1: Comparison of empirical data against the two models across different categories (Authors, CEOs, etc.).
Deep Dive into Results
The findings offer a nuanced view of social influence:
- The Power of Niche: When users followed specific types of celebrities—CEOs, Authors, or Entrepreneurs—the Crandall model fit the data remarkably well (validated via Kolmogorov-Smirnov tests). This suggests that "professional interest" is a highly structured signal for social clustering.
- The Noise of General Fame: When the model looked at the entire celebrity set (Plot f), the Crandall model failed. Why? Because following "Mega-Celebrities" (Pop stars, movie icons) is a ubiquitous behavior that crosses all social boundaries, making it "noisy" and useless for tie inference.
- Upper Bound Consistency: The Kossinets model held true as a universal upper bound, proving that while friendship is probabilistic, it rarely exceeds the theoretical coincidence limit of independent activities.
Figure 2: Probability of friendship as a function of the number of common activities .
Critical Insight & Conclusion
The study proves that homophily (the tendency of individuals to associate with similar others) is quantifiable even through high-level interest data.
Takeaways:
- Categorization Matters: Not all "common activities" are equal. To infer ties accurately, one must weight "exclusive" or "niche" activities more heavily than "popular" ones.
- Privacy Implications: Even if you don't share your friend list, your subscription to technical or niche newsletters/profiles can leak your social circle to an observer.
Limitations:
The researchers noted that the Crandall model, while good for niches, lacked a sufficiently complex "underlying network structure" to handle the chaos of general celebrity following. Future work needs to incorporate social graph topology alongside activity frequency for a more robust prediction engine.
