De-Anonymizing the Social Graph: Privacy Risk Analysis for Facebook Users

Facebook Kullanıcıları İçin Gizlilik Riski Analizi Privacy Risk Analysis for Facebook Users

Önder Çoban, Ali İnan, Selma Özel, Çukurova Üniversitesi, Mühendisligi Bölümü, Türkiye Adana
Summary
Problem
Method
Results
Takeaways
Abstract

The paper presents a privacy risk analysis for Facebook users in Turkey using Intrinsic Privacy Scores (IPS) and Network-aware Privacy Scores (NPS). By crawling real-world data from 10,000 accounts, the authors quantify risk based on profile attribute visibility and network connectivity using PageRank-based algorithms.

TL;DR

This study moves beyond theoretical surveys to analyze the actual privacy exposure of Facebook users in Turkey. By implementing Intrinsic (IPS) and Network-aware (NPS) privacy scores on a real-world dataset of 10,000 users, the research reveals that demographic factors like gender (male) and age (21-40) are primary indicators of high risk. The work highlights that your privacy is not just defined by what you share, but by your position within the social hardware.

Problem & Motivation: The Gap Between Policy and Reality

Online Social Networks (OSNs) offer granular privacy settings, yet a massive 93% of users remain concerned about data access. Why?

  1. Complexity: Users often find settings too convoluted to manage.
  2. Breadth-First Exposure: Information shared with "friends of friends" can be easily harvested by automated crawlers.
  3. The Network Effect: Even if you are private, your presence in a high-visibility friend group can lead to your data being inferred via machine learning.

The authors argue that to fix this, we must first be able to quantify risk using real data rather than subjective questionnaires.

Methodology: Calculating the Cost of Sharing

The research employs two sophisticated metrics to map out the privacy landscape:

1. Intrinsic Privacy Score (IPS)

IPS focuses on the Information Revelation (R-matrix). It calculates sensitivity () for 21 attributes (e.g., address, political views, gender).

  • Sensitivity: A dynamic value—if an attribute is rarely shared in the network, its exposure by a specific user is considered more "sensitive."
  • Visibility: Scaled from 0 (Private) to 3 (Public).

2. Network-aware Privacy Score (NPS)

This is where the math gets interesting. Using the PageRank algorithm, the authors treat "Privacy Risk" like "Web Authority." If you are connected to many high-risk (highly visible) individuals, your own risk score (NPS) increases. It accounts for the structural vulnerability of your social circle.

Privacy Score Transformation Logic Figure 1: Conceptual visualization of how network connections influence individual risk levels (from IPS to NPS).

Experiments & Results: Who is Most at Risk?

The study crawled two datasets (5,000 and 10,000 users). The findings provide a clear demographic profile of privacy vulnerability:

  • The Gender Gap: Male users consistently showed higher risk levels than female users across both datasets.
  • The Age Factor: The 21-40 age group is the most exposed. This "digital native" cohort shares more attributes (education, workplace, location) to facilitate social and professional networking, often at the cost of privacy.
  • Attribute Sensitivity: While Gender and Education are frequently shared, Email and Phone numbers remain the most guarded, significantly affecting the IPS weightings.

User Distribution by Risk Level Table 1: User distribution across NPS levels. Note that while most users have low NPS, a critical mass occupies "Medium" to "High" risk tiers as the network grows.

Critical Analysis & Conclusion

Takeaway

The core contribution of this work is the application of the Pensa et al. framework to a specific, real-world regional context. It proves that automated crawlers can easily reconstruct a user's risk profile using only public-facing data.

Limitations

A significant hurdle identified by the authors is the "Seed Account Bias." Since a crawler sees the network from a specific starting point, it might underestimate the visibility of certain attributes if it isn't "friends" with the target. This suggests the actual privacy risk in the wild is likely higher than the study reports.

Future Outlook

For future OSN development, these scoring mechanisms could be integrated into user dashboards. Imagine a "Privacy Credit Score" that warns you: "Adding these 5 friends will increase your Network Privacy Risk by 20%." This proactive approach is the next frontier in digital self-defense.


Author Bias/Affiliation: The research was conducted by teams from Çukurova University and Adana Alparslan Türkeş Science and Technology University.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use machine learning to predict hidden user attributes on Facebook based on friend list visibility.
  • Which study first introduced the concept of the 'Privacy Score' in social networks, and how has the mathematical definition evolved from linear weighting to graph-based propagation?
  • How are privacy risk scoring models like IPS and NPS being adapted for decentralized social networks or privacy-preserving graph neural networks?
Contents
De-Anonymizing the Social Graph: Privacy Risk Analysis for Facebook Users
1. TL;DR
2. Problem & Motivation: The Gap Between Policy and Reality
3. Methodology: Calculating the Cost of Sharing
3.1. 1. Intrinsic Privacy Score (IPS)
3.2. 2. Network-aware Privacy Score (NPS)
4. Experiments & Results: Who is Most at Risk?
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook