Crowdsourcing and GDPR: Balancing Public Wisdom with Personal Privacy

An assessment of privacy preservation in crowdsourcing approaches: Towards GDPR compliance

2018-05-01
Vasiliki Diamantopoulou, Aggeliki Androutsopoulou, Stefanos Gritzalis, Yannis Charalabidis
Summary
Problem
Method
Results
Takeaways
Abstract

This paper assesses the privacy implications of three crowdsourcing methods in public policy-making—active, passive, and passive expert-sourcing—against GDPR requirements. It identifies significant privacy risks and proposes a mapping between specific privacy requirements and Privacy Enhancing Technologies (PETs) using an adjacency matrix.

TL;DR

As governments increasingly turn to Web 2.0 and social media to crowdsource public policy ideas, they face a looming legal and ethical hurdle: GDPR compliance. This paper evaluates three core crowdsourcing paradigms—Active, Passive, and Passive Expert-sourcing—to identify where they fail privacy standards and provides a roadmap for integrating Privacy Enhancing Technologies (PETs) to fix them.

The Motivation: The Privacy Paradox in Public Participation

Crowdsourcing tools like citizen-sourcing (gathering ideas from the public) and expert-sourcing (leveraging specialist knowledge) are powerful. However, the raw data used—comments, likes, and social media posts—contains significant Personal Identifiable Information (PII).

The authors argue that most existing ICT platforms for e-participation fail to account for the "dark side" of data mining:

  • Profiling: Users can be tracked across different platforms.
  • Identity Theft: Aggregated data can reveal sensitive health or religious views.
  • Legal Liability: Under GDPR, organizations (data controllers) must provide rights like the "Right to Erasure" and "Right to be Forgotten," which are difficult to manage in big data environments.

Methodology: The Three Paradigms of Crowdsourcing

The paper categorizes crowdsourcing into three structural methods, each with different data flows and risks:

  1. Active Crowdsourcing: Government poses a specific problem on its own channels, stimulating discussion.
  2. Passive Crowdsourcing: Government monitors external social media (blogs, Twitter) to find pre-existing policy-related content.
  3. Passive Expert-Sourcing: Automatically retrieving and ranking content from recognized experts.

Data Processing Flow

Table of Data Processed The table above highlights the specific PII collected, from simple metrics like 'likes' to complex documents.

The Core Assessment: Privacy Requirements vs. Reality

The authors evaluated these methods against fundamental privacy requirements. While Authentication (knowing who someone claims to be) and Authorization (controlling access) were generally well-handled by social media APIs, more complex requirements were found wanting.

  • Anonymity: Infringed in expert-sourcing because the expert's reputation is tied to their identity.
  • Unlinkability: A major risk in passive crowdsourcing; a third party could link a specific comment back to a user's profile, leading to surveillance.
  • Undetectability: Currently and often compromised, as attackers can detect if a specific "Item of Interest" (sensitive comment) exists within the crowd data.

Compliance Adjacency Matrix

Adjacency Matrix This matrix provides a "heat map" of compliance. Note the 'X' markers indicating where expert-sourcing intentionally breaks privacy to establish credibility.

Moving Toward Solutions: Privacy Process Patterns

To bridge the gap between legal theory (GDPR) and technical code, the paper proposes Privacy Process Patterns. These are reusable blueprints that implement PETs based on the specific requirement at risk:

  • For Anonymity: Implement Mix-nets, Onion Routing (Tor-like), and Aggregation Gateways to decouple data from its source.
  • For Unlinkability: Use surrogate keys and "Do Not Track" patterns in identity federation.
  • For Undetectability: Deploy steganography and encryption tools for transactions and individual documents.

Conclusion and Future Outlook

The takeaway for policy-makers and developers is clear: Trust is the currency of participation. If citizens feel that sharing their "wisdom" will lead to profiling or state surveillance, crowdsourcing efforts will fail.

The authors conclude that while current systems are technically functional, they are legally vulnerable. The next phase of research must move beyond theoretical mapping toward the implementation of these patterns into the core architecture of e-participation platforms. The goal is a "Privacy by Default" ecosystem where democratic engagement doesn't come at the cost of personal liberty.

Find Similar Papers

Try Our Examples

  • Search for recent studies that implement specific Privacy Enhancing Technologies (PETs) within active crowdsourcing platforms for government policy-making.
  • Which paper first formally defined the 'Privacy By Design' framework, and how does this paper's 'Privacy Process Patterns' expand upon that original theory?
  • Examine how differential privacy and federated learning have been applied to passive crowdsourcing to maintain high-quality sentiment analysis while ensuring GDPR compliance.
Contents
Crowdsourcing and GDPR: Balancing Public Wisdom with Personal Privacy
1. TL;DR
2. The Motivation: The Privacy Paradox in Public Participation
3. Methodology: The Three Paradigms of Crowdsourcing
3.1. Data Processing Flow
4. The Core Assessment: Privacy Requirements vs. Reality
4.1. Compliance Adjacency Matrix
5. Moving Toward Solutions: Privacy Process Patterns
6. Conclusion and Future Outlook