Out of the Wild: Leveraging Contextual Integrity for Default Privacy in Social Ecosystems
Out of the wild: On generating default policies in social ecosystems
This paper introduces a semantic-web-based privacy model for social ecosystems that leverages Helen Nissenbaum's Theory of Contextual Integrity to automatically generate fine-grained default privacy policies. It utilizes OWL ontologies and SPARQL queries to enforce context-aware access control across aggregated social data from disparate platforms like Facebook and LinkedIn.
TL;DR
As we move toward "social ecosystems"—platforms that aggregate data from Facebook, LinkedIn, and personal sensors—privacy management becomes a nightmare. This paper presents a framework to solve the "lazy user" problem (where >90% of people stick to default settings) by using Semantic Web technologies to automatically generate default privacy policies based on the theory of Contextual Integrity. It ensures that your boss stays in the "Professional" context and your gaming buddies stay in the "Gaming" context, even if their data lives in the same database.
The "Default Policy" Crisis
Modern privacy is broken because it relies on manual configuration. Research shows that nearly 99% of Twitter users and 87% of Facebook users never touch their privacy settings. In an aggregated social ecosystem, this is dangerous. If you merge your professional data with your leisure data, a "public" or "friends-of-friends" default policy in one context might accidentally leak sensitive info into another.
The authors argue that the problem is the lack of a formal framing for what constitutes "right and wrong" in data flow.
Methodology: Privacy as Contextual Integrity
The core insight of this work is the adoption of Helen Nissenbaum’s Contextual Integrity (CI). CI posits that privacy isn't about secrecy, but about the appropriate flow of information within specific social contexts.
The Semantic Architecture
The authors propose a layered "Social Hourglass" architecture:
- Social Sensors: Collect data from specific domains (e.g., a LinkedIn sensor for professional data).
- Social Ecosystems Knowledge Base (SEKB): An OWL-based ontology that stores data as triples.
- Privacy Management Layer: The brain of the system, which generates SPARQL queries to act as gates.
Fig 1: The layered architecture separating data collection from privacy enforcement.
Norms of Appropriateness & Distribution
To implement CI, the system defines two types of norms:
- Norms of Appropriateness: Defines what information is okay to share. If Alice is a "Colleague" in Bob's "Professional Context," she cannot request data from his "Gaming Context."
- Norms of Distribution: Defines how it is shared. For example, if a photo is "Shared" or "Co-owned" between Bob and Charlie, the system defaults to requiring Bob’s consent before Charlie can leak it to a third party.
Experimental Implementation: The Aegis Prototype
The researchers implemented a prototype called Aegis using the Jena framework and Java SE 6. They move away from the heavy, centralized reasoning of traditional Semantic Web Rule Language (SWRL) systems.
Instead, they use SPARQL ASK queries. When a request comes in (e.g., "Can Alice see Bob's group?"), the system "augments" the policy with the specific requestor's URI and checks if the triple exists in the knowledge base.
Fig 2: A partial definition of the social ecosystems ontology showing distinct friendship, gaming, and professional contexts.
Critical Analysis & SOTA Comparison
Compared to previous trust-based models (like FOAF-Realm) or relationship-based access control (ReBAC), this approach has several advantages:
- Context-Awareness: It acknowledges that "relationships" are not global. You are a "Friend" only in specific contexts.
- Lower Maintenance: Because it uses SPARQL on localized knowledge bases, it doesn't require re-computing the entire world's permissions whenever one person changes a status.
- Extensibility: New contexts (like "Health" or "Finance") can be added by simply importing new ontologies.
Limitations
While powerful, the model assumes that data can be cleanly assigned to a single context. In the real world, "Professional" and "Friendship" often overlap (e.g., a colleague who is also a close friend). The system's current "disjoint context" assumption may need to evolve into a fuzzy logic or multi-context membership model.
Conclusion
This paper moves privacy from a "manual chore" to an "architectural property." By encoding the social norms of human interaction into the data layer through Semantic Web tools, the authors provide a roadmap for building social ecosystems that protect users by default, even when the users themselves are "out in the wild."
