SELECT * FROM USER: How Database Schemas Script Our Social Lives
SELECT * FROM USER: Infrastructure and Socio-technical Representation
This paper presents a socio-technical analysis of Facebook and craigslist Missed Connections, examining how their underlying computational infrastructures shape digital identity and social interaction. By applying Agre’s eight features of computing practice, the authors demonstrate how rigid database schemas (Facebook) versus transient anonymity (craigslist) fundamentally dictate how users "represent" themselves and their relationships.
TL;DR
At the intersection of computer science and sociology, this paper argues that the code and databases underlying social media are not just tools—they are "representational infrastructures" that dictate the boundaries of our digital existence. By comparing the rigid, persistent world of Facebook with the anonymous, transient world of craigslist Missed Connections, researchers Jed Brubaker and Gillian Hayes reveal how data modeling choices (like the "Friend" entity) force complex human sociality into reductive binary structures.
The Problem: The Socio-Technical Gap
In the world of CSCW (Computer Supported Cooperative Work), there is a fundamental mismatch between what is socially required and what is technically feasible. This is the "Socio-Technical Gap."
Traditional research views social media as a reflection of reality. Brubaker and Hayes challenge this, suggesting that systems like Facebook don't just "store" your life; they "perform" it through a lens of developer-chosen standards. When a system designer decides that a relationship is either "Single" or "In a Relationship," they are not just creating a database field—they are establishing a social ontology that users must then navigate, often with real-world emotional consequences.
Methodology: Peering into the Black Box
The researchers didn't just interview users; they reverse-engineered the systems.
- Facebook Analysis: They analyzed third-party APIs and SDKs to understand the "Ontology" of a person as defined by Facebook (User ID, Friend ID, Network ID).
- craigslist Scrutiny: They collected over 550,000 posts to see how users "work around" a system that has almost no built-in infrastructure for identity.
The Architecture of Identity
The paper utilizes Agre’s framework to dissect these platforms. Two areas stand out:
1. From "Friends" to "Entities" (Ontology & Standards)
In the physical world, friendships are asymmetrical and ambiguous. On Facebook, the database requires a bidirectional confirmed link.

The system imposes a "Standard" that treats your spouse the same as an intern you met once. This homogenization creates "Performance" anxiety—users must curate their profiles for an audience that the database cannot distinguish.
2. The Authenticity of Anonymity (Authentication & Instrumentation)
Craigslist Missed Connections lacks profiles. This creates an authentication problem. How do you know the person responding to your ad is actually the "cute blonde in the navy shirt"? Users developed a social workaround: Custom Instrumentation. They began asking "challenge questions" (e.g., "Tell me what my t-shirt said") to verify identity, effectively building a verification layer that the developers never coded.
The Bias of the Schema
A critical insight of the paper is the inherent bias in selection.
- Gender: Both systems reify the gender binary. Facebook’s original schema forced a choice between "Male" and "Female," excluding transgender or intersex identities from the fundamental data layer.
- Temporality: Facebook's infrastructure values persistence (data that grows forever), while craigslist values transience (data that expires in 7 days). This technical choice changes behavior: Facebook is for "stalking" existing connections; craigslist is for "capturing" lost moments.
Conclusion: Future-Proofing Design
Brubaker and Hayes conclude that we must stop viewing "users" and "systems" as separate. As we move toward more automated social systems:
- Design is never static: Users will always appropriate systems to fit their "worldview," creating workarounds for rigid code.
- Databases are political: The choice of what to "Select" for a database entity (and what to leave out) is a powerful act of social engineering.
The "Socio-Technical Representation" reminds us that when we "SELECT * FROM USER," we aren't just getting data—we are getting a filtered, scripted version of humanity defined by the architecture of the machine.
