Graphing the Identity: Leveraging Social Networks for Author Name Disambiguation

Automatic Method for Author Name Disambiguation Using Social Networks

2010-01-01
Dongwook Shin, Taehwan Kim, Hana Jung, Joongmin Choi
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces an automatic Author Name Disambiguation (AND) method that leverages social networks constructed from bibliographic data (DBLP). By utilizing co-author relationships and semantic topic extraction from paper abstracts, the system identifies both "namesakes" (multiple people with one name) and "heteronymous names" (one person with multiple name variations).

TL;DR

In the academic world, the "John Smith" problem is a nightmare for data integrity. This paper tackles Author Name Disambiguation (AND) by treating bibliographic records as a social network. By detecting structural cycles among co-authors and performing semantic analysis on abstracts, the researchers built a system capable of splitting namesakes and merging heteronymous identities with an F-measure of nearly 88%.

The Identity Crisis in Digital Libraries

Identity uncertainty occurs when a unique identifier is missing for an object. In digital libraries like DBLP, this manifests in two frustrating ways:

  1. Namesakes: Different authors sharing an identical name (e.g., two different "Taehwan Kim"s at different universities).
  2. Heteronymous Names: One author appearing under multiple variations (e.g., "Tom Mitchell" vs. "Tom M. Mitchell").

Traditional methods focusing on string similarity often fail because they ignore the context of collaboration. If two "Tom Mitchells" never share a co-author or a research topic, they are likely different people, even if their names are identical.

Methodology: Beyond Simple Matching

The authors propose a multi-stage architecture to clean bibliographic data:

1. Social Network Construction

The system extracts metadata (authors, titles) and enriches it by crawling the web for abstracts and affiliations. A graph is built where vertices represent authors and edges represent co-authorship.

2. Namesake Detection via Cycle Analysis

A key technical insight is the use of Cycle Detection. If a single vertex (name) acts as a bridge between two otherwise separate, dense clusters of co-authors (cycles), it likely represents two distinct people merged into one node.

  • Algorithm: Based on Johnson’s cycle enumeration.
  • Semantic Check: The system measures the cosine similarity between the "Topic Candidates" (noun phrases from abstracts) of these cycles. If similarity is below a threshold (), the vertex is split.

System Architecture

3. Merging Heteronymous Names

To catch different spellings of the same person, the system uses:

  • LCS (Longest Common Subsequence): To handle abbreviations.
  • Organization Matching: Verifying if "T. Mitchell" and "Tom Mitchell" belong to the same institution.
  • Sharing Vertex Detection: If two similar names share a common co-author, they are likely the same person.

Experimental Results

The researchers tested their approach on DBLP data. The results show a clear progression in performance as each module (Namesake Splitter, OrgMatcher, SharingVertex) is activated.

Performance Comparison

Key Finding: Using only basic social network analysis (SNA) yielded an F-measure of 83.1% under standard metrics. However, after applying the full pipeline including namesake detection and sharing vertex detection, the system became much more robust, achieving a high recall of 94.3%.

The authors also introduced an "Improved Evaluation Method" to penalize "partial matches," arguing that in data integration, the quality of the match matters as much as the quantity. Under this stricter regime, the system maintained a respectable 71% F-measure, proving its utility in real-world, "noisy" environments.

Critical Analysis & Future Outlook

This paper's strength lies in its hybrid approach: it combines the topological structure of social graphs with the semantic richness of NLP (noun phrase extraction).

Limitations:

  • Dynamic Affiliations: The "Organization Matcher" might struggle with researchers who move frequently between institutions.
  • Computational Cost: Cycle detection in massive graphs (millions of authors) can be computationally expensive without heuristic pruning.

The Path Forward:

The authors suggest that future work will involve refining the cycle detection algorithms to handle even larger, noisier datasets. In the context of 2026, this logic is a precursor to modern Graph Neural Networks (GNNs) which now automate much of this feature engineering.

Takeaway: Effective disambiguation isn't just about how a name is spelled; it's about the "company" the author keeps.

Find Similar Papers

Try Our Examples

  • Search for recent papers that use Graph Neural Networks (GNNs) or Deep Walk embeddings to solve the author name disambiguation problem in DBLP or PubMed datasets.
  • Which 1975 paper by Donald B. Johnson first defined the elementary circuit enumeration algorithm used for cycle detection in this study?
  • Examine how social network-based name disambiguation methods have been adapted for entity linking in heterogeneous knowledge graphs or multi-modal social media data.
Contents
Graphing the Identity: Leveraging Social Networks for Author Name Disambiguation
1. TL;DR
2. The Identity Crisis in Digital Libraries
3. Methodology: Beyond Simple Matching
3.1. 1. Social Network Construction
3.2. 2. Namesake Detection via Cycle Analysis
3.3. 3. Merging Heteronymous Names
4. Experimental Results
5. Critical Analysis & Future Outlook
5.1. Limitations:
5.2. The Path Forward: