Disambiguating the Web: A Multi-Factor Framework for Person Name Resolution
Person Name Disambiguation in Web Pages Using Social Network, Compound Words and Latent Topics
The paper introduces a multi-pronged framework for Web-based person name disambiguation, integrating Social Network extraction, Compound Word statistics, and Latent Topic modeling. By combining these three distinct feature sets, the authors achieve a state-of-the-art F-measure of 0.7529 on a Japanese Web document dataset.
TL;DR
When you search for "John Smith," the web returns a chaotic mix of CEOs, athletes, and local teachers. This paper presents a robust framework to automatically cluster these pages by individual. By fusing Social Network (SN) analysis, Compound Key Word (CKW) statistics, and Dirichlet Process Unigram Mixture (DPUM) models, the authors achieve a significant performance leap (F-measure ~0.75) over traditional word-frequency baselines.
Context & Motivation
Identity ambiguity is a fundamental challenge in information retrieval. Existing solutions often struggle because:
- Links are sparse: Direct hyperlinks between pages of the same person are rarer than one might think.
- Names are culturally varied: Many systems are hard-coded for Western "First Middle Last" name structures, failing on Japanese or other formats.
- The number of entities is unknown: Unlike standard classification, we don't know in advance how many "John Smiths" exist in the dataset.
The authors' core intuition is that an individual is defined by who they know (Social Network), how they talk (Compound Words), and what they talk about (Latent Topics).
Methodology: The Three Pillars
The framework processes documents through three distinct pipelines:
1. Social Network (SN) Extraction
The system uses a Named Entity (NE) tagger to identify other people, places, and organizations co-occurring with the target name.
- Logic: If two pages mention the same "rare" colleague's name alongside the target, they likely refer to the same person.
- Weighting: Person names are assigned higher weights () than places or organizations ().
2. Compound Key Words (CKW)
Instead of single words, the system extracts compound nouns (e.g., "Information Technology Center").
- Why?: Compound words are far more discriminative than single nouns.
- The Metric: They use a specialized scoring formula that considers the frequency of the compound word and the variety of neighboring nouns (LR score), capturing the "terminological status" of the phrase.
3. Latent Topic Modeling (DPUM)
While traditional topic models like LDA require you to pre-set the number of topics (), this paper employs a Dirichlet Process.
- The "Stick-Breaking" Intuition: Imagine a stick of unit length representing the total probability. The process repeatedly breaks off pieces to assign to new topics. This allows the model to "grow" the number of entities (clusters) based on the data itself, which is perfect for an unknown number of people.
Fig 1: Conceptual visualization of Social Network linkages between documents.
Experimental Insights & Results
The authors tested their framework on a manually annotated dataset of 5,015 Japanese Web pages across 38 names.
| Method | F-Measure | Precision | Recall |
|---|---|---|---|
| Baseline (Word Freq) | 0.5409 | 0.6668 | 0.6950 |
| SN (Social Network) | 0.7163 | 0.9000 | 0.6692 |
| CKW (Compound Words) | 0.6974 | 0.8195 | 0.7050 |
| (SN ∪ CKW) + DP | 0.7529 | 0.8496 | 0.7640 |
Key Takeaways from Experiments:
- Complementarity: SN provides high precision (it's hard to share a social network by accident), while CKW and DPUM help improve recall by capturing broader content patterns.
- The Power of "Union": Combining SN and CKW via a Union (connecting pages if either method matches) yielded a 4-5 point jump over using them individually.
- DP Refinement: Initializing the Dirichlet Process with clusters from SN/CKW filters out noise and refines the semantic boundaries of each person's cluster.
Comparison of F-measures showing the consistent advantage of the combined approach (SN ∪ CKW) across different person names.
Critical Analysis & Conclusion
This work successfully moves beyond simple bag-of-words clustering. By using non-parametric Bayes (DPUM), it elegantly solves the problem of not knowing how many "entities" are in a search result.
Limitations:
- The system assumes "one page = one entity," which fails on compilation pages (like alumni lists or news aggregate pages).
- It relies heavily on the quality of the Named Entity Tagger; if the NE tagger misses a key colleague's name, the Social Network link breaks.
Future Outlook: With the rise of Large Language Models (LLMs), the "Latent Topic" part of this framework could potentially be replaced by dense vector embeddings. However, the explicit use of Social Networks remains a highly interpretable and powerful "anchor" for identity that even modern black-box models benefit from.
