SELECT * FROM USER: How Database Schemas Script Our Social Lives

SELECT * FROM USER: Infrastructure and Socio-technical Representation

2015-10-29
Jed R. Brubaker, Gillian R. Hayes
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a socio-technical analysis of Facebook and craigslist Missed Connections, examining how their underlying computational infrastructures shape digital identity and social interaction. By applying Agre’s eight features of computing practice, the authors demonstrate how rigid database schemas (Facebook) versus transient anonymity (craigslist) fundamentally dictate how users "represent" themselves and their relationships.

TL;DR

At the intersection of computer science and sociology, this paper argues that the code and databases underlying social media are not just tools—they are "representational infrastructures" that dictate the boundaries of our digital existence. By comparing the rigid, persistent world of Facebook with the anonymous, transient world of craigslist Missed Connections, researchers Jed Brubaker and Gillian Hayes reveal how data modeling choices (like the "Friend" entity) force complex human sociality into reductive binary structures.

The Problem: The Socio-Technical Gap

In the world of CSCW (Computer Supported Cooperative Work), there is a fundamental mismatch between what is socially required and what is technically feasible. This is the "Socio-Technical Gap."

Traditional research views social media as a reflection of reality. Brubaker and Hayes challenge this, suggesting that systems like Facebook don't just "store" your life; they "perform" it through a lens of developer-chosen standards. When a system designer decides that a relationship is either "Single" or "In a Relationship," they are not just creating a database field—they are establishing a social ontology that users must then navigate, often with real-world emotional consequences.

Methodology: Peering into the Black Box

The researchers didn't just interview users; they reverse-engineered the systems.

  • Facebook Analysis: They analyzed third-party APIs and SDKs to understand the "Ontology" of a person as defined by Facebook (User ID, Friend ID, Network ID).
  • craigslist Scrutiny: They collected over 550,000 posts to see how users "work around" a system that has almost no built-in infrastructure for identity.

The Architecture of Identity

The paper utilizes Agre’s framework to dissect these platforms. Two areas stand out:

1. From "Friends" to "Entities" (Ontology & Standards)

In the physical world, friendships are asymmetrical and ambiguous. On Facebook, the database requires a bidirectional confirmed link.

General Infrastructure and Context

The system imposes a "Standard" that treats your spouse the same as an intern you met once. This homogenization creates "Performance" anxiety—users must curate their profiles for an audience that the database cannot distinguish.

2. The Authenticity of Anonymity (Authentication & Instrumentation)

Craigslist Missed Connections lacks profiles. This creates an authentication problem. How do you know the person responding to your ad is actually the "cute blonde in the navy shirt"? Users developed a social workaround: Custom Instrumentation. They began asking "challenge questions" (e.g., "Tell me what my t-shirt said") to verify identity, effectively building a verification layer that the developers never coded.

The Bias of the Schema

A critical insight of the paper is the inherent bias in selection.

  • Gender: Both systems reify the gender binary. Facebook’s original schema forced a choice between "Male" and "Female," excluding transgender or intersex identities from the fundamental data layer.
  • Temporality: Facebook's infrastructure values persistence (data that grows forever), while craigslist values transience (data that expires in 7 days). This technical choice changes behavior: Facebook is for "stalking" existing connections; craigslist is for "capturing" lost moments.

Conclusion: Future-Proofing Design

Brubaker and Hayes conclude that we must stop viewing "users" and "systems" as separate. As we move toward more automated social systems:

  1. Design is never static: Users will always appropriate systems to fit their "worldview," creating workarounds for rigid code.
  2. Databases are political: The choice of what to "Select" for a database entity (and what to leave out) is a powerful act of social engineering.

The "Socio-Technical Representation" reminds us that when we "SELECT * FROM USER," we aren't just getting data—we are getting a filtered, scripted version of humanity defined by the architecture of the machine.

Find Similar Papers

Try Our Examples

  • Search for recent CSCW or CHI papers that discuss the 'socio-technical gap' in modern algorithmic social media feeds.
  • Which paper by Philip Agre first conceptualized the 'eight features of computing practice,' and how has it been applied to modern AI-driven platforms?
  • Find studies examining how non-binary gender identities are represented or excluded in the database schemas of contemporary social networking sites.
Contents
SELECT * FROM USER: How Database Schemas Script Our Social Lives
1. TL;DR
2. The Problem: The Socio-Technical Gap
3. Methodology: Peering into the Black Box
4. The Architecture of Identity
4.1. 1. From "Friends" to "Entities" (Ontology & Standards)
4.2. 2. The Authenticity of Anonymity (Authentication & Instrumentation)
5. The Bias of the Schema
6. Conclusion: Future-Proofing Design