Beyond Friendships and Followers: Mapping the Global Social Fabric through Wikipedia

2015 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining 472

Summary
Problem
Method
Results
Takeaways
Abstract

The paper presents a framework for constructing a large-scale, person-centric social network from the unstructured text of English Wikipedia. By leveraging Wikidata and interwiki links, it identifies ~800k persons and ~67M edges, achieving a state-of-the-art representation of global historical and contemporary social structures.

TL;DR

Researchers from Heidelberg University have moved beyond the "friend" button to map the social architecture of human history. By mining over 5.6 million Wikipedia articles and linking them to Wikidata, they constructed the Wikipedia Social Network (WSN)—a massive graph of 800,000 people. Unlike Twitter or Facebook, this network is built on shared context and narrative proximity, offering a unique lens into how historical and contemporary figures are interconnected.

The Problem: The "Explicit Bias" of Social Data

Most social network analysis (SNA) today is tethered to "explicit" networks like LinkedIn or Twitter. While rich, these datasets are limited: they only record people who are currently online, they are often proprietary, and they are shallow in historical depth.

Previous attempts to extract "latent" networks from text (like co-authorship in DBLP) were too niche. The challenge was creating a network that was massive, diverse, and semantically accurate without drowning in the noise of accidental name mentions.

Methodology: From Text to Topology

The researchers' insight was to treat Wikipedia not just as a collection of facts, but as a "Latent Space" of human relationships.

1. Identifying the Nodes (The Wikidata Backbone)

Rather than relying on error-prone Named Entity Recognition (NER), the authors used Interwiki Links (IWLs). By cross-referencing Wikipedia links with Wikidata IDs, they ensured nearly 100% precision in identifying persons, ranging from Barack Obama to Napoleon.

2. The Weight of Distance: The "Dicos" Metric

The core innovation lies in how they defined a "relationship." If two people are mentioned in the same article, are they related? Maybe. If they are in the same sentence? Almost certainly. They introduced an exponential decay function for edge weights: where is the number of sentences between mentions. This "distance-weighted cosine similarity" (dicos) allows the network to distinguish between a "List of Basketball Players" (weakly related) and a "Political Rivalry" (strongly related).

Overall Distribution of Links Fig 1: The distribution of links across Wikipedia articles shows a long-tail pattern typical of high-quality information repositories.

Experiments & Results: A Mirror of Reality

The WSN behaves remarkably like a real-world social network.

  • Assortativity: People in the network tend to connect with others of similar "importance" (degree).
  • Temporal Logic: The probability of a link drops sharply if the time span between two people’s birth dates exceeds a human lifespan, validating the "contextual" relevance of the mentions.
  • Centrality: Using PageRank, the authors identified global hubs. Unsurprisingly, figures like Barack Obama, John Paul II, and Napoleon topped the list, reflecting their cross-domain influence.

PageRank Centrality Table Table 1: Top 20 most central figures in the Wikipedia Social Network.

Community Detection: Discovering the "Uncategorized"

Using the Stabilized Label Propagation Algorithm (SLPA), the team identified thousands of communities. Interestingly, these communities often outperformed Wikipedia’s own internal "Category" system. For instance, the network successfully grouped the entire rowing crew of the 1962 British Empire Games—a group so specific it lacks a dedicated Wikipedia category.

Critical Insight: Why This Matters

The Wikipedia Social Network serves as a "ground truth" for human history. While Facebook shows us who we know, the WSN shows us who matters in the collective narrative of humanity.

Limitations:

  • Gender Bias: 84.3% of the nodes are male, reflecting the historical and systemic biases present in Wikipedia’s content.
  • NLP Constraints: The study primarily focused on the English Wikipedia, potentially missing regional social hubs in other languages.

Future Outlook: By integrating this graph with modern LLMs, we could potentially automate person name disambiguation (e.g., which "John Smith" is this article about?) by looking at the social neighborhood of the mention. It turns Wikipedia from a book of facts into a living map of human interaction.

Find Similar Papers

Try Our Examples

  • Which recent papers have utilized the Wikipedia Social Network or similar Wikidata-derived graphs for large-scale entity disambiguation or link prediction tasks?
  • Beyond simple co-occurrence, what are the current SOTA methods for extracting typed relationships (e.g., "spouse of", "rival of") from unstructured biographical text using LLMs or Relation Extraction (RE) models?
  • How does the community structure of person-centric networks in Wikipedia compare to the graph topology of modern explicit social networks like Twitter or Mastodon in terms of information diffusion patterns?
Contents
Beyond Friendships and Followers: Mapping the Global Social Fabric through Wikipedia
1. TL;DR
2. The Problem: The "Explicit Bias" of Social Data
3. Methodology: From Text to Topology
3.1. 1. Identifying the Nodes (The Wikidata Backbone)
3.2. 2. The Weight of Distance: The "Dicos" Metric
4. Experiments & Results: A Mirror of Reality
4.1. Community Detection: Discovering the "Uncategorized"
5. Critical Insight: Why This Matters