Crowdsourcing and GDPR: Balancing Public Wisdom with Personal Privacy
An assessment of privacy preservation in crowdsourcing approaches: Towards GDPR compliance
This paper assesses the privacy implications of three crowdsourcing methods in public policy-making—active, passive, and passive expert-sourcing—against GDPR requirements. It identifies significant privacy risks and proposes a mapping between specific privacy requirements and Privacy Enhancing Technologies (PETs) using an adjacency matrix.
TL;DR
As governments increasingly turn to Web 2.0 and social media to crowdsource public policy ideas, they face a looming legal and ethical hurdle: GDPR compliance. This paper evaluates three core crowdsourcing paradigms—Active, Passive, and Passive Expert-sourcing—to identify where they fail privacy standards and provides a roadmap for integrating Privacy Enhancing Technologies (PETs) to fix them.
The Motivation: The Privacy Paradox in Public Participation
Crowdsourcing tools like citizen-sourcing (gathering ideas from the public) and expert-sourcing (leveraging specialist knowledge) are powerful. However, the raw data used—comments, likes, and social media posts—contains significant Personal Identifiable Information (PII).
The authors argue that most existing ICT platforms for e-participation fail to account for the "dark side" of data mining:
- Profiling: Users can be tracked across different platforms.
- Identity Theft: Aggregated data can reveal sensitive health or religious views.
- Legal Liability: Under GDPR, organizations (data controllers) must provide rights like the "Right to Erasure" and "Right to be Forgotten," which are difficult to manage in big data environments.
Methodology: The Three Paradigms of Crowdsourcing
The paper categorizes crowdsourcing into three structural methods, each with different data flows and risks:
- Active Crowdsourcing: Government poses a specific problem on its own channels, stimulating discussion.
- Passive Crowdsourcing: Government monitors external social media (blogs, Twitter) to find pre-existing policy-related content.
- Passive Expert-Sourcing: Automatically retrieving and ranking content from recognized experts.
Data Processing Flow
The table above highlights the specific PII collected, from simple metrics like 'likes' to complex documents.
The Core Assessment: Privacy Requirements vs. Reality
The authors evaluated these methods against fundamental privacy requirements. While Authentication (knowing who someone claims to be) and Authorization (controlling access) were generally well-handled by social media APIs, more complex requirements were found wanting.
- Anonymity: Infringed in expert-sourcing because the expert's reputation is tied to their identity.
- Unlinkability: A major risk in passive crowdsourcing; a third party could link a specific comment back to a user's profile, leading to surveillance.
- Undetectability: Currently and often compromised, as attackers can detect if a specific "Item of Interest" (sensitive comment) exists within the crowd data.
Compliance Adjacency Matrix
This matrix provides a "heat map" of compliance. Note the 'X' markers indicating where expert-sourcing intentionally breaks privacy to establish credibility.
Moving Toward Solutions: Privacy Process Patterns
To bridge the gap between legal theory (GDPR) and technical code, the paper proposes Privacy Process Patterns. These are reusable blueprints that implement PETs based on the specific requirement at risk:
- For Anonymity: Implement Mix-nets, Onion Routing (Tor-like), and Aggregation Gateways to decouple data from its source.
- For Unlinkability: Use surrogate keys and "Do Not Track" patterns in identity federation.
- For Undetectability: Deploy steganography and encryption tools for transactions and individual documents.
Conclusion and Future Outlook
The takeaway for policy-makers and developers is clear: Trust is the currency of participation. If citizens feel that sharing their "wisdom" will lead to profiling or state surveillance, crowdsourcing efforts will fail.
The authors conclude that while current systems are technically functional, they are legally vulnerable. The next phase of research must move beyond theoretical mapping toward the implementation of these patterns into the core architecture of e-participation platforms. The goal is a "Privacy by Default" ecosystem where democratic engagement doesn't come at the cost of personal liberty.
