Reality Mining as an Engineering Discipline: Predicting the Limits of Social Sensing

Incremental Learning with Accuracy Prediction of Social and Individual Properties from Mobile-Phone Data

2012-09-01
Yaniv Altshuler, Nadav Aharony, Michael Fire, Yuval Elovici, Alex Pentland
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a structured methodology for incremental learning from mobile-phone data, leveraging the "Friends and Family" dataset from MIT. It utilizes classifiers like Random Forest and Naive-Bayes to predict attributes such as ethnicity, student status, and life partners, while proposing a Gompertz-function-based model to predict the limits of learning accuracy.

TL;DR

Researchers from the MIT Media Lab have moved mobile data analysis from a "craft" to a "science." By analyzing the "Friends and Family" dataset, they've demonstrated that the accuracy of predicting personal traits (ethnicity, occupation) and social ties (life partners) follows a predictable mathematical growth curve known as the Gompertz function. This allows scientists to predict the maximum possible accuracy of a model using only the first few days of data collection.

Background: Beyond the "Artisan" Data Scientist

For years, "Reality Mining"—the extraction of human behavior from sensor data—has been a field of trial and error. Data scientists relied on "gut feelings" to decide how long to collect data or which signals to prioritize. This is inefficient, especially considering the power constraints of mobile phones. The core question this paper answers is: Can we mathematically predict the value of future data before we even collect it?

Methodology: The Math of Learning

The authors utilized the "Friends and Family" dataset, containing 20 million WiFi scans and over 200,000 calls from 140 participants. They focused on several classification tasks:

  • Ethnicity: Using the Louvain method for community detection on SMS networks.
  • Student Status: Using Rotation-Forest classifiers.
  • Significant Others: Analyzing Bluetooth collocation social graphs.

The breakthrough was modeling the accuracy () over time () using the Gompertz function: This function, commonly used to describe tumor growth or technology adoption, perfectly captures the "diminishing returns" of data collection.

Model Architecture: SMS and Bluetooth Social Networks Figure 1: SMS Social Network Graph where node colors represent ethnicity clusters identified by the Louvain method.

Key Results: Achieving Saturation

The findings show that prediction accuracy doesn't grow linearly; it accelerates and then levels off. For example, in the "Significant Other" classification task, the model achieved a 65.6% success rate, with the Gompertz regression providing a near-perfect fit for the learning trajectory.

Accuracy Evolution and Regression Figure 2: The prediction accuracy of the Ethnicity classifier over time, modeled by the Gompertz function.

Crucially, the authors demonstrated that by extrapolating this function, researchers can decide to stop data collection once the "saturation point" is reached, saving battery life and computational resources.

Learning Process Extrapolation Figure 3: Extrapolation of learning curves in linear and log-log scales, illustrating the ability to predict future performance.

Critical Insight: Data Correlations

The paper also explores how the learning dynamics of different traits are correlated. High correlation between learning curves (e.g., Origin vs. Significant Other) suggests that these features are linked in the real world—confirming sociological observations like homophily (people marrying within their own ethnic group).

From a technical perspective, this means if one signal is "expensive" (like high-frequency GPS) and another is "cheap" (call logs), and their learning curves are correlated, we might skip the expensive signal entirely.

Conclusion

This research provides a building block for the "Science of Big Data." By treating accuracy as a predictable growth process, the authors offer a toolset for designing more efficient, less intrusive, and mathematically grounded social sensing systems. The major limitation remains the "phone-owner proxy" assumption—the model assumes owners and phones are always together, which isn't always true, potentially explaining the accuracy ceiling.

Takeaway for Practitioners

Don't just collect data indefinitely. Use the first 7-10 days of your pilot study to fit a growth curve; if the projected "saturation" accuracy is below your requirements, you likely need different features, not more of the same data.

Find Similar Papers

Try Our Examples

  • Search for recent papers that apply the Gompertz function or other sigmoidal growth models to predict machine learning performance as a function of dataset size.
  • Which paper first introduced the "Friends and Family" dataset at MIT, and what were the primary sensors used for initial behavioral sensing?
  • Find studies that compare mobile-based social sensing accuracy between Bluetooth proximity data and GPS-based location proximity for predicting romantic relationships.
Contents
Reality Mining as an Engineering Discipline: Predicting the Limits of Social Sensing
1. TL;DR
2. Background: Beyond the "Artisan" Data Scientist
3. Methodology: The Math of Learning
4. Key Results: Achieving Saturation
5. Critical Insight: Data Correlations
6. Conclusion
6.1. Takeaway for Practitioners