Disambiguating the Web: A Multi-Factor Framework for Person Name Resolution

Person Name Disambiguation in Web Pages Using Social Network, Compound Words and Latent Topics

2008-05-10
Shingo Ono, Issei Sato, Minoru Yoshida, Hiroshi Nakagawa
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a multi-pronged framework for Web-based person name disambiguation, integrating Social Network extraction, Compound Word statistics, and Latent Topic modeling. By combining these three distinct feature sets, the authors achieve a state-of-the-art F-measure of 0.7529 on a Japanese Web document dataset.

TL;DR

When you search for "John Smith," the web returns a chaotic mix of CEOs, athletes, and local teachers. This paper presents a robust framework to automatically cluster these pages by individual. By fusing Social Network (SN) analysis, Compound Key Word (CKW) statistics, and Dirichlet Process Unigram Mixture (DPUM) models, the authors achieve a significant performance leap (F-measure ~0.75) over traditional word-frequency baselines.

Context & Motivation

Identity ambiguity is a fundamental challenge in information retrieval. Existing solutions often struggle because:

  • Links are sparse: Direct hyperlinks between pages of the same person are rarer than one might think.
  • Names are culturally varied: Many systems are hard-coded for Western "First Middle Last" name structures, failing on Japanese or other formats.
  • The number of entities is unknown: Unlike standard classification, we don't know in advance how many "John Smiths" exist in the dataset.

The authors' core intuition is that an individual is defined by who they know (Social Network), how they talk (Compound Words), and what they talk about (Latent Topics).

Methodology: The Three Pillars

The framework processes documents through three distinct pipelines:

1. Social Network (SN) Extraction

The system uses a Named Entity (NE) tagger to identify other people, places, and organizations co-occurring with the target name.

  • Logic: If two pages mention the same "rare" colleague's name alongside the target, they likely refer to the same person.
  • Weighting: Person names are assigned higher weights () than places or organizations ().

2. Compound Key Words (CKW)

Instead of single words, the system extracts compound nouns (e.g., "Information Technology Center").

  • Why?: Compound words are far more discriminative than single nouns.
  • The Metric: They use a specialized scoring formula that considers the frequency of the compound word and the variety of neighboring nouns (LR score), capturing the "terminological status" of the phrase.

3. Latent Topic Modeling (DPUM)

While traditional topic models like LDA require you to pre-set the number of topics (), this paper employs a Dirichlet Process.

  • The "Stick-Breaking" Intuition: Imagine a stick of unit length representing the total probability. The process repeatedly breaks off pieces to assign to new topics. This allows the model to "grow" the number of entities (clusters) based on the data itself, which is perfect for an unknown number of people.

Architecture Logic Fig 1: Conceptual visualization of Social Network linkages between documents.

Experimental Insights & Results

The authors tested their framework on a manually annotated dataset of 5,015 Japanese Web pages across 38 names.

MethodF-MeasurePrecisionRecall
Baseline (Word Freq)0.54090.66680.6950
SN (Social Network)0.71630.90000.6692
CKW (Compound Words)0.69740.81950.7050
(SN ∪ CKW) + DP0.75290.84960.7640

Key Takeaways from Experiments:

  1. Complementarity: SN provides high precision (it's hard to share a social network by accident), while CKW and DPUM help improve recall by capturing broader content patterns.
  2. The Power of "Union": Combining SN and CKW via a Union (connecting pages if either method matches) yielded a 4-5 point jump over using them individually.
  3. DP Refinement: Initializing the Dirichlet Process with clusters from SN/CKW filters out noise and refines the semantic boundaries of each person's cluster.

Performance across different queries Comparison of F-measures showing the consistent advantage of the combined approach (SN ∪ CKW) across different person names.

Critical Analysis & Conclusion

This work successfully moves beyond simple bag-of-words clustering. By using non-parametric Bayes (DPUM), it elegantly solves the problem of not knowing how many "entities" are in a search result.

Limitations:

  • The system assumes "one page = one entity," which fails on compilation pages (like alumni lists or news aggregate pages).
  • It relies heavily on the quality of the Named Entity Tagger; if the NE tagger misses a key colleague's name, the Social Network link breaks.

Future Outlook: With the rise of Large Language Models (LLMs), the "Latent Topic" part of this framework could potentially be replaced by dense vector embeddings. However, the explicit use of Social Networks remains a highly interpretable and powerful "anchor" for identity that even modern black-box models benefit from.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize non-parametric Bayesian models like Dirichlet Process for entity resolution in large-scale Web search results.
  • Which paper first established the "stick-breaking process" for Dirichlet Process mixtures, and how has it been optimized for high-dimensional text clustering since then?
  • Are there recent studies that apply the combination of social network graphs and latent topic models to cross-document coreference resolution in multi-modal contexts?
Contents
Disambiguating the Web: A Multi-Factor Framework for Person Name Resolution
1. TL;DR
2. Context & Motivation
3. Methodology: The Three Pillars
3.1. 1. Social Network (SN) Extraction
3.2. 2. Compound Key Words (CKW)
3.3. 3. Latent Topic Modeling (DPUM)
4. Experimental Insights & Results
4.1. Key Takeaways from Experiments:
5. Critical Analysis & Conclusion