Mining Microdata: How Machine Learning Settles a 200-Year-Old Debate on Social Mobility

Mining microdata: Economic opportunity and spatial mobility in Britain and the United States, 1850–1881

2014-10-01
Peter Baskerville, Lisa Dillon, Kris Inwood, Evan Roberts, Steven Ruggles, Kevin Schurer, John Robert Warren, Peter Baskerville, Lisa Dillon, Kris Inwood, Evan Roberts, Steven Ruggles, Kevin Schurer, John Robert Warren
Summary
Problem
Method
Results
Takeaways
Abstract

This paper investigates intergenerational social mobility in Britain and the United States between 1850 and 1881 by applying Support Vector Machine (SVM) and probabilistic record linkage to census microdata. The authors utilize the North Atlantic Population Project (NAPP) database to construct longitudinal panels, ultimately finding that the United States exhibited significantly higher economic and geographic mobility than Great Britain.

TL;DR

For centuries, historians have debated whether 19th-century America was truly a "land of opportunity" compared to class-bound Europe. By applying Support Vector Machines (SVM) and probabilistic record linkage to massive census microdata from 1850–1881, this research proves that the U.S. was significantly more mobile than Great Britain. While British sons were often "locked" into their father's social strata, American sons—particularly those of unskilled laborers—found much higher rates of upward movement, often facilitated by westward geographic migration.

The "Safety Valve" vs. The "Floating Proletariat"

Since Alexis de Tocqueville visited North America in the 1830s, the "equality of conditions" has been a central pillar of American identity. However, 20th-century revisionist historians argued this was a myth, suggesting that high geographic mobility actually masked a "floating proletariat" of workers who moved frequently but stayed poor.

The core difficulty in resolving this debate has been data scale. Linking an individual across censuses taken 30 years apart is a "needle in a haystack" problem. Traditional manual linking is slow and prone to bias, while exact-string matching fails to account for 19th-century spelling inconsistencies and transcription errors.

Methodology: High-Precision Record Linkage

The authors utilized the North Atlantic Population Project (NAPP), creating longitudinal panels through a sophisticated machine-learning pipeline.

1. Avoiding Selection Bias

Most commercial record linkage (like genealogy software) uses "stable" markers like a spouse’s name or city of residence. The authors intentionally ignored these. Why? Because using a spouse's name biases the sample toward those who stayed married, and using location biases it toward those who didn't move—the very variables they were trying to measure.

2. The SVM Pipeline

The team used a Support Vector Machine (SVM) to classify potential matches. To handle the "John Smith" problem—where multiple people shard the same name and age—the algorithm:

  • Used Jaro-Winkler string similarity to handle minor spelling errors.
  • Applied Phonetic Encoding (NYSIIS and Double-Metaphone) to group names that sound the same but are spelled differently.
  • Set a strict confidence threshold where a link was only established if one and only one candidate exceeded the requirement.

需替换为架构图 Note: The project methodology relied on blocking factors (sex, race, birthplace) to reduce 15 trillion potential comparisons in the 1880 U.S. census to a computationally feasible scale.

Key Findings: The Atlantic Divide

The results confirm a stark difference in the social fabric of the two nations.

Geographic Restlessness

In the U.S., 52% of men moved counties between 1850 and 1880, with a clear westward trend. In Britain, only 36% moved, and these were typically short-distance moves to adjacent counties.

Occupational Fluidity

The occupational inheritance—sons doing exactly what their fathers did—was much stronger in Britain.

  • Unskilled Workers: In Britain, 44% of sons of unskilled workers remained unskilled. In the U.S., that figure dropped to 27%, with the rest moving into skilled labor, farming, or white-collar roles.
  • The Altham Statistic: To compare these two different economic structures (one heavily agrarian, one industrial), the authors used the Altham statistic (). This mathematical tool measures the "distance from independence." The results showed that the link between father and son was 66% stronger in Britain than in the U.S.

实验结果对比 Table IV: Comparison of occupational mobility matrices. Note the higher churn out of the "Unskilled" row in the U.S. data compared to Great Britain.

Deep Insight: The Value of Land

A unique portion of the study analyzed acreage in British farming. The authors found that a son in Britain had a less than 50% chance of remaining a farmer unless his father owned more than 400 acres. In contrast, the easy availability of land in the U.S. acted as an "escape valve," allowing sons of even the poorest laborers to acquire property and change their social standing.

Critical Analysis & Future Outlook

While the study provides robust evidence for 19th-century American mobility, it faces a few limitations:

  • Gender Bias: The linkage relies heavily on surnames, meaning women (who typically changed names upon marriage) are excluded from the current panel.
  • The "Farmer" Catch-all: The U.S. census often labeled everyone in agriculture simply as a "Farmer," potentially hiding nuances in wealth and status within that broad category.

Conclusion: This work demonstrates that "Big Data" history, powered by machine learning, can provide quantitative clarity to qualitative debates. It suggests that institutional and environmental factors—like land policy and the timing of the Industrial Revolution—fundamentally shaped the disparate "life chances" of citizens on either side of the Atlantic. The next step? Incorporating Canadian data to see if the "North American effect" holds true across the entire continent.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use Deep Learning or Transformer-based models for historical census record linkage to compare performance against SVM-based probabilistic methods.
  • Which seminal paper first introduced the Altham statistic for comparing contingency tables in social sciences, and how has it been adapted for long-term historical mobility studies?
  • Explore longitudinal studies that apply the Historical International Standard Classification of Occupations (HISCO) to compare 19th-century social mobility in mainland Europe versus North America.
Contents
Mining Microdata: How Machine Learning Settles a 200-Year-Old Debate on Social Mobility
1. TL;DR
2. The "Safety Valve" vs. The "Floating Proletariat"
3. Methodology: High-Precision Record Linkage
3.1. 1. Avoiding Selection Bias
3.2. 2. The SVM Pipeline
4. Key Findings: The Atlantic Divide
4.1. Geographic Restlessness
4.2. Occupational Fluidity
5. Deep Insight: The Value of Land
6. Critical Analysis & Future Outlook