Mining the Past: How Social Network Properties Reveal Name Inconsistencies in DBLP

Learning from the Past: An Analysis of Person Name Corrections in DBLP Collection and Social Network Properties of Affected Entities

2010-08-01
Florian Reitz, Oliver Hoffmann
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a longitudinal study of name-related data quality in the DBLP digital library by mining person name corrections over a ten-year period. It leverages a "historic DBLP collection" to categorize errors—such as synonyms and homonyms—and analyzes their distribution within dynamic social networks of coauthorship and publication streams.

TL;DR

Researchers at the University of Trier have analyzed ten years of historical data from the DBLP bibliographic project to understand why name errors occur. By tracking how "defective" names were corrected over time, they discovered that an author's position in a social network—who they coauthor with and where they publish—is a strong predictor of whether their name record is accurate or requires correction.

Background: The Identity Crisis in Digital Libraries

Digital libraries like DBLP are the backbone of academic research, but they face a persistent challenge: Name Ambiguity. This manifests in two primary ways:

  • Homonyms: Multiple people sharing the same name (e.g., the 19 "Wei Wangs" in DBLP).
  • Synonyms: One person being listed under multiple names due to spelling errors, transcription differences, or name changes.

While automated algorithms exist, they are rarely perfect. The gold standard remains manual correction, often triggered by "community feedback" from authors themselves. This paper asks: Can we learn from these past human-made corrections to predict future errors?

Methodology: The Historic DBLP Framework

The authors reconstructed the history of DBLP by mining 3,300 backups from 1999 to 2009. They defined four types of modifications to track how data quality improved over time:

  1. Rename: A simple label change (e.g., "H. Schweppe" to "Heinz Schweppe").
  2. Merge: Combining two names into one (resolving synonyms).
  3. Split: Dividing one name into two (resolving homonyms).
  4. Distribute: Reassigning a specific paper from one author to another.

Analyzing Network DNA

The research evaluated these entities through two distinct social networks:

  • Collaboration Network (C): A graph where nodes are authors and edges represent coauthorship.
  • Name-Stream Network (S): A two-mode network linking authors to the "streams" (conferences or journals) where they publish.

Model Architecture Fig 1: Conceptual mapping between real-world persons and DBLP name entities.

Key Insights: Error Profiles

The study found significant differences in how "incorrect" names exist within the network:

  • Local Connectivity: Names that were eventually split or merged tended to have a very low Clustering Coefficient. Essentially, if your "coauthors" don't know each other (not forming a clique), you might actually be two different people mistakenly merged into one record.
  • Global Centrality: Interestingly, defective entities often showed high Betweenness Centrality. This suggests that prominent authors with high publication counts are more likely to have errors—partly because they have more opportunities for data entry mistakes, and partly because the community monitors their records more closely.

Graph of Degree Distribution Fig 2: Distribution of node degrees in the collaboration network, showing how split/distribute candidates often have larger neighborhoods than stable nodes.

A New Heuristic: Assessing "Reliable Areas"

The authors proposed that errors aren't just random—they cluster. Some publishers or conferences (streams) have "noisier" metadata than others. By calculating a Stream Node Risk (rS) and a Collaboration Risk (rC), they can assign a reliability score to any given name.

In their evaluation:

  • Safe Zones: Name entities in the top-tier "Reliable Areas" had very low correction rates.
  • Danger Zones: Entities in the bottom-tier "Error Prone Areas" (often containing Spanish or Chinese names with incomplete first names) were 3x more likely to be corrected in the following two years.

Critical Analysis & Future Outlook

This work provides a unique "Ground Truth" for the research community. While most papers try to solve disambiguation with the current state of a database, this paper proves that the evolutionary history of the database is a goldmine for improving data quality.

Limitations: The study is biased toward errors that DBLP’s current tools and community are good at catching. Errors that remain undiscovered by humans are, by definition, missing from this analysis.

Takeaway: For builders of knowledge graphs and digital libraries, the message is clear: don't just look at the node—look at the neighborhood. If an author publishes in streams with high historical error rates and has a "star-shaped" coauthor network, their record likely requires a closer look.

Find Similar Papers

Try Our Examples

  • Look for recent papers that utilize graph neural networks (GNNs) or social network embeddings for author name disambiguation in the DBLP dataset.
  • Which study first defined the standard taxonomy of name disambiguation errors (synonyms/homonyms) in bibliographic databases, and how does this paper's categorization expand upon it?
  • Investigate how modern persistent identifiers like ORCID have impacted the rate of manual name corrections in digital libraries compared to the period analyzed in this study.
Contents
Mining the Past: How Social Network Properties Reveal Name Inconsistencies in DBLP
1. TL;DR
2. Background: The Identity Crisis in Digital Libraries
3. Methodology: The Historic DBLP Framework
3.1. Analyzing Network DNA
4. Key Insights: Error Profiles
5. A New Heuristic: Assessing "Reliable Areas"
6. Critical Analysis & Future Outlook