From Mentions to Kinship: Building Family Networks through Bootstrapping and CDCR

Research on Building Family Networks Based on Bootstrapping and Coreference Resolution

2013-01-01
Jinghang Gu, Yanan Hu, Longhua Qian, Qiaoming Zhu
Summary
Problem
Method
Results
Takeaways
Abstract

The paper proposes a novel framework for building personal family networks from large-scale Chinese corpora using bootstrapping and cross-document coreference resolution (CDCR). By iteratively learning relational patterns and instances, it extracts "Parent-Child" and "Husband-Wife" relationships and aggregates them into multi-person family structures while addressing name ambiguity and variation.

TL;DR

Researchers from Soochow University have developed a pipeline to automatically construct family networks from massive text corpora (Gigaword). By combining a minimally supervised bootstrapping approach (Espresso) with robust Cross-Document Coreference Resolution (CDCR), the system can identify "who is who" and "who belongs to whom" across millions of news articles, achieving a relation precision of 94%.

The Core Challenge: Beyond Simple Co-occurrence

Most social network analysis tools rely on simple name co-occurrence: if "Alice" and "Bob" appear in the same sentence, they are linked. However, this fails in two major ways:

  1. Ambiguity: Is "Simone" in Document A (Kahn's wife) the same "Simone" in Document B (Gbagbo's wife)?
  2. Missing Context: It treats individuals as isolated nodes, ignoring the foundational unit of human society—the family.

Methodology: The Three-Step Pipeline

1. Bootstrapping Relation Extraction

The authors define two primary relation types: Parent-Child and Husband-Wife. Using an iterative process, the system starts with "seed" pairs (e.g., Jiang Zemin and Wang Yeping).

  • Pattern Induction: Finding strings between names (e.g., "...'s wife...").
  • Reliability Scoring: Using Pointwise Mutual Information (PMI) to ensure patterns like "X and Y" (too generic) are filtered out, while specific patterns are promoted.

Need to replace with Figure/Formula of Reliability Calculation

2. Solving the Identity Crisis (CDCR)

Once relations are extracted, the system must merge them. This is where most systems fail. The authors use a dual-strategy:

  • Name Disambiguation (ND): Uses the "context" of a name (other entities in the document) to create a feature vector. If the cosine similarity between two "Simones" is low, they are treated as different people.
  • Name Variation Aggregation (NVA): Uses Levenshtein (edit) distance to recognize that "Suharto" and "Suharduo" might be the same person despite spelling differences.

3. Family Network Aggregation

The final step clusters entities. A valid "family" requires at least three persons and two relationships. This prevents noise from accidental links.

Experimental Insights

The study evaluated the system on the Gigaword corpus (over 1 million articles).

Relation Extraction Performance

Relation TypePatterns FoundExamples
Parent-Child26<Parent>’s son <Child>, <Parent>’s daughter <Child>
Husband-Wife33<Husband>’s wife <Wife>, <Husband>’s widow <Wife>

The precision for relation extraction reached 94.0%, proving that bootstrapping is highly effective even with minimal manual seeds.

The CDCR Impact

The researchers found that Name Variation (spelling differences) is a much bigger problem than Name Ambiguity in news text.

  • Using Exact Name Matching (ENM) only gave an F1-score of 45.0 for family fusion.
  • Adding Name Variation Aggregation (NVA) boosted the F1-score to 59.4.
  • The best performance (60.3 F1) came from combining Disambiguation and Variation Aggregation.

Table of Results Placeholder

Critical Analysis & Conclusion

Takeaway

This research shifts the focus of Social Network Analysis from "nodes" to "units." The success of the CDCR phase proves that linguistic variations are the primary bottleneck in building accurate social graphs.

Limitations

  • Recall Issues: The system still misses families where name expressions are extremely diverse.
  • Limited Relations: Focuses only on immediate family. It does not yet map complex extended families (cousins, in-laws) or inter-family alliances.

Future Outlook

The authors suggest that the next frontier is mapping the associations between different families, potentially leading to a "Global Social Map" that reflects the actual structure of human society more accurately than current keyword-based graphs.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Large Language Models (LLMs) to solve the Cross-Document Coreference Resolution (CDCR) problem in personal relation extraction.
  • Which paper first introduced the Espresso bootstrapping algorithm, and how does the pattern reliability scoring in this work differ from the original formulation?
  • Explore research that applies family network extraction or kinship discovery to genealogical reconstruction or historical document analysis.
Contents
From Mentions to Kinship: Building Family Networks through Bootstrapping and CDCR
1. TL;DR
2. The Core Challenge: Beyond Simple Co-occurrence
3. Methodology: The Three-Step Pipeline
3.1. 1. Bootstrapping Relation Extraction
3.2. 2. Solving the Identity Crisis (CDCR)
3.3. 3. Family Network Aggregation
4. Experimental Insights
4.1. Relation Extraction Performance
4.2. The CDCR Impact
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook