From Mentions to Kinship: Building Family Networks through Bootstrapping and CDCR
Research on Building Family Networks Based on Bootstrapping and Coreference Resolution
The paper proposes a novel framework for building personal family networks from large-scale Chinese corpora using bootstrapping and cross-document coreference resolution (CDCR). By iteratively learning relational patterns and instances, it extracts "Parent-Child" and "Husband-Wife" relationships and aggregates them into multi-person family structures while addressing name ambiguity and variation.
TL;DR
Researchers from Soochow University have developed a pipeline to automatically construct family networks from massive text corpora (Gigaword). By combining a minimally supervised bootstrapping approach (Espresso) with robust Cross-Document Coreference Resolution (CDCR), the system can identify "who is who" and "who belongs to whom" across millions of news articles, achieving a relation precision of 94%.
The Core Challenge: Beyond Simple Co-occurrence
Most social network analysis tools rely on simple name co-occurrence: if "Alice" and "Bob" appear in the same sentence, they are linked. However, this fails in two major ways:
- Ambiguity: Is "Simone" in Document A (Kahn's wife) the same "Simone" in Document B (Gbagbo's wife)?
- Missing Context: It treats individuals as isolated nodes, ignoring the foundational unit of human society—the family.
Methodology: The Three-Step Pipeline
1. Bootstrapping Relation Extraction
The authors define two primary relation types: Parent-Child and Husband-Wife. Using an iterative process, the system starts with "seed" pairs (e.g., Jiang Zemin and Wang Yeping).
- Pattern Induction: Finding strings between names (e.g., "...'s wife...").
- Reliability Scoring: Using Pointwise Mutual Information (PMI) to ensure patterns like "X and Y" (too generic) are filtered out, while specific patterns are promoted.

2. Solving the Identity Crisis (CDCR)
Once relations are extracted, the system must merge them. This is where most systems fail. The authors use a dual-strategy:
- Name Disambiguation (ND): Uses the "context" of a name (other entities in the document) to create a feature vector. If the cosine similarity between two "Simones" is low, they are treated as different people.
- Name Variation Aggregation (NVA): Uses Levenshtein (edit) distance to recognize that "Suharto" and "Suharduo" might be the same person despite spelling differences.
3. Family Network Aggregation
The final step clusters entities. A valid "family" requires at least three persons and two relationships. This prevents noise from accidental links.
Experimental Insights
The study evaluated the system on the Gigaword corpus (over 1 million articles).
Relation Extraction Performance
| Relation Type | Patterns Found | Examples |
|---|---|---|
| Parent-Child | 26 | <Parent>’s son <Child>, <Parent>’s daughter <Child> |
| Husband-Wife | 33 | <Husband>’s wife <Wife>, <Husband>’s widow <Wife> |
The precision for relation extraction reached 94.0%, proving that bootstrapping is highly effective even with minimal manual seeds.
The CDCR Impact
The researchers found that Name Variation (spelling differences) is a much bigger problem than Name Ambiguity in news text.
- Using Exact Name Matching (ENM) only gave an F1-score of 45.0 for family fusion.
- Adding Name Variation Aggregation (NVA) boosted the F1-score to 59.4.
- The best performance (60.3 F1) came from combining Disambiguation and Variation Aggregation.

Critical Analysis & Conclusion
Takeaway
This research shifts the focus of Social Network Analysis from "nodes" to "units." The success of the CDCR phase proves that linguistic variations are the primary bottleneck in building accurate social graphs.
Limitations
- Recall Issues: The system still misses families where name expressions are extremely diverse.
- Limited Relations: Focuses only on immediate family. It does not yet map complex extended families (cousins, in-laws) or inter-family alliances.
Future Outlook
The authors suggest that the next frontier is mapping the associations between different families, potentially leading to a "Global Social Map" that reflects the actual structure of human society more accurately than current keyword-based graphs.
