From Paper Cards to Graph Databases: Digitizing the WWII Japanese-American Incarceration Records

Digital Curation of a World War II Japanese-American Incarceration Camp Collection: Implications for Sociotechnical Archival Systems

2018-10-01
Richard Marciano, Myeong Lee, William Underwood, Sandra Laib, Zeynep Diker, Aakanksha Singh
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a Computational Archival Science (CAS) framework applied to the digital curation of 25,000 WWII Japanese-American Incarceration Camp index cards. It utilizes a custom NLP/NER pipeline within the DRAS-TIC repository to automate item-level metadata extraction, social network analysis, and privacy-preserving redaction.

TL;DR

This research presents a transformative approach to archival science by applying Computational Archival Science (CAS) to a massive collection of WWII Japanese-American Incarceration Camp records. By moving beyond simple scanning, the team leverages NLP pipelines, social network analysis, and scalable repositories to turn 25,000 index cards into a searchable, relational, and ethically curated digital ecosystem.

The Archive Crisis: Scale and Silence

For decades, hundreds of millions of item-level records—from WWI award cards to WWII internment rosters—have remained locked in administrative limbo. The problem isn't just that they are analog; it’s that traditional cataloging methods cannot handle the volume.

The War Relocation Authority (WRA) records are particularly sensitive. They contain "internal security" reports on Japanese-Americans imprisoned during WWII. These aren't just names; they are stories of "disorderly conduct, theft, and accidents" recorded by camp staff. Simply digitizing them as images leaves the real history hidden.

The Methodology: Archiving as a Sociotechnical System

The authors argue that archiving in the digital age is not just a technical challenge but a sociotechnical one. Their workflow doesn't just rely on code; it relies on the interaction between algorithms and human expertise.

1. The NLP/NER Pipeline

Using the GATE (General Architecture for Text Engineering) software and the ANNIE plugin, the team developed a workflow to extract more than just names. They target:

  • Case Report IDs
  • Housing Identifiers
  • Offense Types
  • Organizations and Locations

Model Architecture: The Multi-Faceted Tagging Process

2. DRAS-TIC: Scaling the Repository

To manage this metadata, they utilized DRAS-TIC (Digital Repository At Scale That Invites Computation). Unlike traditional static repositories, DRAS-TIC uses REST APIs and NoSQL back-ends to allow near real-time searching of extracted NLP terms, effectively turning a static image into a queryable data point.

Unlocking Social Data

One of the most profound insights of this research is the use of Social Network Analysis (SNA). By linking cards that share names, dates, and incidents, the researchers can visualize the hidden social structures within the camps.

Social Network Visualization of Overlapping Camp Incidents

This graph-based modeling allows researchers to see how individuals were interconnected across different incident contexts, providing a much richer historical narrative than a simple list of names.

The Privacy Challenge: Protecting Internment History

Digitization often clashes with privacy. The National Archives (NARA) mandates that information regarding internees who were 18 or younger at the time of the incident must not be released.

To solve this, the researchers built Name Gazetteers using historical data files to determine the ages of individuals on the cards automatically. This "Privacy Mining" ensures that the digital release of the archive respects both historical transparency and individual ethics.

Critical Analysis & Future Outlook

The value of this work lies in its holistic framework. It acknowledges that:

  1. Algorithms have biases: Government-produced content is inherently biased; community involvement (e.g., Densho.org) is necessary to correct the narrative.
  2. Archival Operations must be preserved: The process of how an OCR error was corrected or how a metadata rule was applied is itself a historical record that belongs in an Archival Information Package (AIP).

Limitation: The current system relies heavily on the quality of OCR. While crowdsourcing helps, the sheer diversity of card styles across different camps remains a significant hurdle for universal automation.

Takeaway: This project serves as a blueprint for the future of "Big Data" in the humanities. It proves that by treating archives as dynamic sociotechnical systems, we can finally unlock the hundreds of millions of records currently gathering dust in national repositories.

Find Similar Papers

Try Our Examples

  • Search for recent studies that integrate Named Entity Recognition (NER) with Graph Databases for large-scale historical archival visualization.
  • Identify the seminal papers on Computational Archival Science (CAS) and how the concept of 'Sociotechnical Systems' has evolved within digital humanities.
  • Explore how automated redaction and PII detection algorithms are being applied to historical government records to comply with contemporary privacy laws.
Contents
From Paper Cards to Graph Databases: Digitizing the WWII Japanese-American Incarceration Records
1. TL;DR
2. The Archive Crisis: Scale and Silence
3. The Methodology: Archiving as a Sociotechnical System
3.1. 1. The NLP/NER Pipeline
3.2. 2. DRAS-TIC: Scaling the Repository
4. Unlocking Social Data
5. The Privacy Challenge: Protecting Internment History
6. Critical Analysis & Future Outlook