From Logs to Links: Mining Professional Social Networks and Expertise
Extracting and Utilizing Social Networks from Log Files of Shared Workspaces
The paper presents a framework for extracting object-centric and user-centric social networks from online shared workspace log files. Using Semantic Web technologies (RDF and SPARQL), it introduces the "Cooperation Index" to quantify collaboration and leverages document history for automated expertise profiling.
TL;DR
This research transforms passive log files from shared workspaces (like BSCW or SharePoint) into active social intelligence. By mapping user actions to documents into a User-Centric Social Network, the authors introduce the Cooperation Index—a mathematical weight that defines how closely people actually work together.
Executive Summary
In many organizations, knowing "who knows what" and "who works with whom" is a matter of guesswork. This paper argues that the "fingerprints" we leave while reading, revising, and creating documents are the most reliable indicators of social structure. By utilizing Semantic Web (RDF) and SPARQL, the authors provide a scalable way to extract expertise and measure collaboration strength directly from the system's history.
Problem & Motivation: The Hidden Pulse of Collaboration
Previous attempts to map social networks in the workplace often failed because they required manual input (e.g., "tag your skills"). People are busy; they don't update their profiles.
The authors identified that log files—the boring records of every "read" or "edit" event—are actually a goldmine. However, the raw data is "Object-Centric" (focused on the file). The challenge was: how do we pivot this data to become "User-Centric" and assign meaningful weights to these relationships?
Methodology: The Object-to-User Pivot
The core innovation lies in the transition from document clouds to weighted human connections.
1. The RDF Mapping
Every log entry is converted into an RDF Triple: [User] -> [Event] -> [Object]. This allows the researchers to use SPARQL (a query language) to find all people who touched the same document.
2. From Object-Centric to User-Centric

The authors "collapse" the document in the middle. If Alice created a file and Bob revised it, a direct link is drawn between Alice and Bob.
3. The Cooperation Index
Not all actions are equal. The researchers propose a weighted sum:
- Create + Revise = High Weight (Deep collaboration)
- Read + Read = Low Weight (Weak connection)
This Cooperation Index quantifies the "collaborative distance" between any two employees based on their shared history.
Prototypes: Holmes and Expert Finder
The authors developed two primary tools to demonstrate the utility of this data:
- Expert Finder: Automatically assigns "expertise elements" to users based on the types of documents they interact with most.
- Holmes: Visualizes the user-centric network and calculates the Cooperation Index between any two selected users.

Experiments & Results
The approach was tested on the Ecospace project involving 183 users. Key findings included:
- Accuracy: 12/12 evaluated participants agreed that the extracted social graph accurately reflected their real-world collaboration.
- Activity Inference: By simply counting RDF triples, the system could identify "power users" who acted as information hubs.
- Dynamic Nature: Unlike static Org Charts, these networks evolve in real-time as project logs grow.
Critical Insight: Why This Matters
The value of this work lies in its Inductive Bias: the assumption that professional relationships are mediated through artifacts (documents). While it might miss "water-cooler talk," it captures the "work" in "workflow."
Limitations: The current model primarily looks at "Depth 1" (people working on the same document). Future advancements could look at "Depth 2" (Alice works with Bob who works with Charlie) to find hidden bridges between isolated departments.
Conclusion: By treating logs not as trash but as a "social sensor," organizations can move toward Hyper-contextual Expertise Finding. No more searching for a "Java expert" in a database; instead, the system tells you who has actually shipped the most code in the last three months.
