Graphing the Identity: Leveraging Social Networks for Author Name Disambiguation
Automatic Method for Author Name Disambiguation Using Social Networks
This paper introduces an automatic Author Name Disambiguation (AND) method that leverages social networks constructed from bibliographic data (DBLP). By utilizing co-author relationships and semantic topic extraction from paper abstracts, the system identifies both "namesakes" (multiple people with one name) and "heteronymous names" (one person with multiple name variations).
TL;DR
In the academic world, the "John Smith" problem is a nightmare for data integrity. This paper tackles Author Name Disambiguation (AND) by treating bibliographic records as a social network. By detecting structural cycles among co-authors and performing semantic analysis on abstracts, the researchers built a system capable of splitting namesakes and merging heteronymous identities with an F-measure of nearly 88%.
The Identity Crisis in Digital Libraries
Identity uncertainty occurs when a unique identifier is missing for an object. In digital libraries like DBLP, this manifests in two frustrating ways:
- Namesakes: Different authors sharing an identical name (e.g., two different "Taehwan Kim"s at different universities).
- Heteronymous Names: One author appearing under multiple variations (e.g., "Tom Mitchell" vs. "Tom M. Mitchell").
Traditional methods focusing on string similarity often fail because they ignore the context of collaboration. If two "Tom Mitchells" never share a co-author or a research topic, they are likely different people, even if their names are identical.
Methodology: Beyond Simple Matching
The authors propose a multi-stage architecture to clean bibliographic data:
1. Social Network Construction
The system extracts metadata (authors, titles) and enriches it by crawling the web for abstracts and affiliations. A graph is built where vertices represent authors and edges represent co-authorship.
2. Namesake Detection via Cycle Analysis
A key technical insight is the use of Cycle Detection. If a single vertex (name) acts as a bridge between two otherwise separate, dense clusters of co-authors (cycles), it likely represents two distinct people merged into one node.
- Algorithm: Based on Johnson’s cycle enumeration.
- Semantic Check: The system measures the cosine similarity between the "Topic Candidates" (noun phrases from abstracts) of these cycles. If similarity is below a threshold (), the vertex is split.

3. Merging Heteronymous Names
To catch different spellings of the same person, the system uses:
- LCS (Longest Common Subsequence): To handle abbreviations.
- Organization Matching: Verifying if "T. Mitchell" and "Tom Mitchell" belong to the same institution.
- Sharing Vertex Detection: If two similar names share a common co-author, they are likely the same person.
Experimental Results
The researchers tested their approach on DBLP data. The results show a clear progression in performance as each module (Namesake Splitter, OrgMatcher, SharingVertex) is activated.

Key Finding: Using only basic social network analysis (SNA) yielded an F-measure of 83.1% under standard metrics. However, after applying the full pipeline including namesake detection and sharing vertex detection, the system became much more robust, achieving a high recall of 94.3%.
The authors also introduced an "Improved Evaluation Method" to penalize "partial matches," arguing that in data integration, the quality of the match matters as much as the quantity. Under this stricter regime, the system maintained a respectable 71% F-measure, proving its utility in real-world, "noisy" environments.
Critical Analysis & Future Outlook
This paper's strength lies in its hybrid approach: it combines the topological structure of social graphs with the semantic richness of NLP (noun phrase extraction).
Limitations:
- Dynamic Affiliations: The "Organization Matcher" might struggle with researchers who move frequently between institutions.
- Computational Cost: Cycle detection in massive graphs (millions of authors) can be computationally expensive without heuristic pruning.
The Path Forward:
The authors suggest that future work will involve refining the cycle detection algorithms to handle even larger, noisier datasets. In the context of 2026, this logic is a precursor to modern Graph Neural Networks (GNNs) which now automate much of this feature engineering.
Takeaway: Effective disambiguation isn't just about how a name is spelled; it's about the "company" the author keeps.
