ArnetMiner: Building Semantic Academic Networks via Unified Extraction and Disambiguation

Social Network Extraction of Academic Researchers

2007-10-01
Jie Tang, Duo Zhang, Limin Yao
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents a comprehensive framework for extracting academic social networks by automatically profiling researchers and integrating publication data. It introduces a unified tagging approach using Conditional Random Fields (CRF) for profile extraction and a constraint-based Hidden Markov Random Field (HMRF) model to solve the name disambiguation problem during data fusion.

TL;DR

Building a professional social network from the Web is a "data-messy" challenge. This paper details the technical backbone of ArnetMiner, a system that transforms unstructured Web pages and scattered publication records into a structured, semantic academic network. It solves two critical bottlenecks: profile extraction via a unified CRF model that understands field dependencies, and name disambiguation using a constraint-based probabilistic model.

Problem & Motivation: The "Silo" Trap

Most information extraction (IE) systems treat different attributes—like a researcher's email, position, and university—as independent variables. However, in reality, these fields are highly correlated. If a text segment identifies "Electrical Engineering," the likelihood that a nearby entity is a "University of Technology" increases significantly. Separate rules or classifiers miss this internal logic.

Furthermore, the identity crisis in academia is rampant. "Jing Zhang" might refer to dozens of different researchers. Standard clustering often fails because it's purely unsupervised, while supervised models don't scale when you have millions of names to clarify. The authors sought a "Middle Way": a semi-supervised approach that can ingest domain knowledge (like co-authorship graphs) as constraints.

Methodology: The Core Architectures

1. Unified Profiling with CRF

Instead of building a dozen separate classifiers, the authors treat a researcher's homepage as a sequence. By using Conditional Random Fields (CRF), the system labels tokens (Standard words, Special words, Images, etc.) based not just on their own features, but also on the labels of surrounding tokens.

  • Visual Logic: Below is the conceptual workflow where tokens from a homepage are mapped into a structured profile schema. Model Architecture: Extraction to Integration

2. Name Disambiguation via HMRF Constraints

To merge DBLP publications with extracted profiles, the authors use a Hidden Markov Random Field (HMRF). This model is unique because it combines a distance metric (how similar are two papers?) with a constraint set.

The system uses six types of constraints:

  • Co-Organization: Do they belong to the same lab?
  • Co-Author: Do they share collaborators?
  • -CoAuthor: A multi-hop co-authorship expansion (e.g., a friend of a friend).
  • Email/Feedback: Hard constraints from unique identifiers or manual corrections.

Experiments & Results: The Power of Context

The researchers proved that "context is king." In the profiling task, the Unified Model (with transition features) significantly outperformed SVM and Amilcare (Rule induction).

Profiling Performance Comparison

Key Findings:

  • Dependency Gains: Removing transition features (Unified_NT) caused an average F1-score drop of ~11%, proving that modeling the relationship between "University" and "Major" is essential.
  • Disambiguation Accuracy: On abbreviated datasets (e.g., "C. Chang"), the proposed method outperformed hierarchical clustering baselines by approximately 8% in F1-score.
  • Real-world Impact: When these techniques were applied to an Expert Finding system, the precision improved dramatically (MAP +22%).

Ablation: Contribution of Features

Deep Insight & Conclusion

This work highlights a pivotal shift in academic knowledge mining: moving from ad-hoc extraction to unified probabilistic modeling. By treating the Web page as a structured sequence and the co-authorship network as a graph of constraints, ArnetMiner successfully handles the intrinsic noise of the internet.

Limitations: The system still relies heavily on the availability of homepages (found for ~70% of researchers). For those without a digital footprint beyond publications, the profile remains sparse. Future work likely involves leveraging heterogeneous graphs to infer missing "latent" profile data.

Takeaway: In any domain where data fields are deeply intertwined—be it medical records, legal documents, or academic profiles—unified sequence labeling (like CRF) and constraint-based clustering (like HMRF) are far superior to "divided and conquer" strategies.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize Deep Learning (like Transformers or GNNs) to solve the author name disambiguation problem in academic databases.
  • Which paper first proposed the use of Hidden Markov Random Fields (HMRF) for semi-supervised clustering, and how does this paper adapt that framework for social network extraction?
  • Explore how contemporary expert finding systems integrate multi-modal data sources (e.g., social media and code repositories) beyond traditional academic homepages and bibliographies.
Contents
ArnetMiner: Building Semantic Academic Networks via Unified Extraction and Disambiguation
1. TL;DR
2. Problem & Motivation: The "Silo" Trap
3. Methodology: The Core Architectures
3.1. 1. Unified Profiling with CRF
3.2. 2. Name Disambiguation via HMRF Constraints
4. Experiments & Results: The Power of Context
5. Deep Insight & Conclusion