PIA: Enhancing Infection Analysis through Privacy-Preserving Social Data Fusion

Exploiting Social Network to Enhance Human-to-Human Infection Analysis without Privacy Leakage

2016-11-08
Kuan Zhang, Xiaohui Liang, Jianbing Ni, Kan Yang, Xuemin Sherman Shen
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces PIA, a Privacy-preserving Infection Analysis approach that integrates social network data with health data to enhance human-to-human infection tracking. By utilizing conditional oblivious transfer and homomorphic encryption, it allows authorized hospitals to analyze infection risks without exposing sensitive user identities or health metrics to untrusted cloud servers.

TL;DR

Predicting how a virus spreads between individuals usually requires a trade-off: either protect privacy and lose accuracy, or track everything and expose sensitive personal data. This paper introduces PIA, a system that combines wearable health monitoring with social network contact data. By using advanced cryptography (Homomorphic Encryption and Oblivious Transfer), it allows health authorities to identify high-risk individuals with 88% accuracy while keeping the underlying data completely encrypted from the cloud servers.

Problem & Motivation: The Context Gap

Traditional epidemic control is "dumb"—it relies on mass quarantines because we lack precise data on who actually had risky contact. While modern wearables can track your temperature (Health Data), they don't know who you sat next to on the bus or for how long (Social Data).

Current solutions fail because:

  1. Lack of Context: Health data alone doesn't show the transmission path.
  2. Privacy Silos: Social networks (like Facebook/WeChat) and Hospitals use independent clouds. Merging their data usually exposes a user's entire social circle and medical history to third-party providers.

The authors' core insight is that infection risk isn't just a biological value; it's a function of Infectivity (Source) × Duration (Contact) × Immunity (Receiver).

Methodology: The Core Engine

The PIA system is built on two primary technical pillars designed to handle data across a "Health Domain" and a "Social Domain."

1. Privacy-Preserving Data Query (PPDQ)

When a doctor identifies a patient as "Infected," they need to find everyone that person recently contacted. To prevent the Social Cloud (SC) from knowing which person is being queried (which would leak their infection status), the system uses Conditional Oblivious Transfer. This allows the hospital to pull specific contact records without the SC learning the identity of the target or the results of the query.

2. Privacy-Preserving Classification (PCIA)

Once the contact list is obtained, the system must calculate the risk. It uses a Naive Bayesian Classifier.

  • The Challenge: How do you run "ArgMax" and "Comparison" logic on data you can't see?
  • The Solution: The authors employ Ring Learning With Error (RLWE) based homomorphic encryption. This allows the Health Cloud to perform additions and multiplications on ciphertexts. They specifically developed a Privacy-Preserving Comparison (PPC) algorithm to handle the non-linear logic required for classification without ever decrypting the data.

System Architecture Figure: The PIA System Model showing the interaction between Trusted Authority, Social Cloud, and Health Cloud.

Experiments & Results

The authors validated PIA using the Infocom06 data set, which contains real Bluetooth-based contact traces of 78 users at a conference.

  • Accuracy Boost: By including social factors (contact duration, social-tie strength) alongside health data, the accuracy of identifying "Susceptible" individuals jumped from 59.16% (baseline) to 85.83%.
  • Computational Feasibility: While homomorphic encryption is notoriously slow, PIA's optimized comparison algorithms keep the Hospital's overhead low (~7 seconds for a batch classification), shifting the heavier the 24-second workload to the powerful Health Cloud.

Performance Comparison Table: Accuracy Comparison highlighting the massive gain (21.5% overall) when leveraging social network data.

Critical Analysis & Conclusion

Takeaway

PIA proves that social networking data is not just for marketing—it is a critical "missing link" in public health. By using Proxy Re-encryption, the system allows users to control their data, only granting decryption rights to the hospital after a diagnosis is confirmed.

Limitations

  • Energy Consumption: The paper notes that continuous Bluetooth/NFC discovery is required on smartphones, which may significantly impact battery life.
  • Trust Model: The system assumes the Hospital is "semi-trusted." In a real-world scenario, a rogue hospital could still potentially abuse its decryption keys if not properly audited.

Future Outlook

The authors suggest moving toward Deep Learning for the next iteration. Integrating Neural Networks with Homomorphic Encryption (NeuroCrypt) would allow for more complex infection patterns (like indirect spread via surfaces) to be analyzed while maintaining the same gold-standard privacy.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Homomorphic Encryption or Secure Multi-Party Computation for infectious disease contact tracing following the COVID-19 pandemic.
  • Which original studies established the "Conditional Oblivious Transfer" protocol, and how does this paper adapt it for social network query privacy?
  • Investigate how more advanced machine learning models, such as Federated Learning or Differential Privacy-based Graph Neural Networks, are being applied to social-health data fusion for epidemic modeling.
Contents
PIA: Enhancing Infection Analysis through Privacy-Preserving Social Data Fusion
1. TL;DR
2. Problem & Motivation: The Context Gap
3. Methodology: The Core Engine
3.1. 1. Privacy-Preserving Data Query (PPDQ)
3.2. 2. Privacy-Preserving Classification (PCIA)
4. Experiments & Results
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook