Predicting Project Outcome from Developer Collaboration Networks: A Discriminative Rich Graph Mining Approach
Predicting Project Outcome Leveraging Socio-Technical Network Patterns
2013-03-01
Summary
Problem
Method
Results
Takeaways
Abstract
This paper introduces the novel problem of predicting software project success/failure by mining discriminative patterns from socio-technical networks represented as rich graphs (multiple labels on nodes/edges). The authors propose a translation method to convert rich graphs into simple graphs, enabling the use of existing discriminative subgraph mining (Top-K LEAP), and prove the translation is sound and complete. Applied to SourceForge.Net data, their approach achieves 94.99% accuracy and 0.86 AUC under 10-fold cross-validation.
## TL;DR
This paper presents a novel framework to predict whether an open-source project will succeed or fail by analyzing the socio-technical network of its developers. The key innovation is a **translation method** that converts a multi-labeled "rich graph" into a "simple graph" that can be processed by existing discriminative subgraph mining algorithms (Top-K LEAP). The mined patterns are then used as features in an SVM classifier, achieving **94.99% accuracy** and **0.86 AUC** on SourceForge.Net data. The work is both sound (all mined patterns are discriminative) and complete (no discriminative pattern is missed).
## Background & Motivation
Every day thousands of software projects are started, but only a fraction achieve widespread adoption. Understanding *why* some projects succeed while others fail has been a long-standing goal in software engineering. Prior work has explored factors like team size, developer experience, and communication patterns, but often relies on manually designed features that may miss complex structural patterns in developer collaboration.
**Key Insight:** The collaboration history among developers forms a *socio-technical network* – a graph where nodes are developers and edges represent past co-work. Each node/edge can carry multiple attributes (e.g., number of past successful projects, length of collaboration). The authors hypothesize that frequent structural patterns in these networks that appear disproportionately in successful vs. failed projects could serve as powerful predictors.
**Challenge:** Existing discriminative subgraph mining algorithms (like LEAP) only work on *simple graphs* where each node and edge has a single label. But socio-technical graphs are *rich graphs* with multiple labels per node/edge. Direct application is impossible.
## Methodology: Translating Rich Graphs to Simple Graphs
The core technical contribution is a sound and complete translation procedure that converts any rich graph into an equivalent simple graph without losing discriminative information.
**Step 1: Model each project as a rich graph**
- Nodes: developers
- Node labels (3): PSP (past successful projects), PFP (past failed projects), LOM (length of membership)
- Edges: collaboration between two developers
- Edge labels (3): PSC (past successful collaborations), PFC (past failed collaborations), LCH (length of collaboration history)
**Step 2: Translate rich graph to simple graph**
The translation involves two types of replicas:
1. **NL-Replicas**: For each node with multiple labels, create one replica per label. Replicas of the same original node are connected by *sibling-replicated edges* (SREs). Edge labels are copied to all replicas.
2. **EL-Replicas**: For each edge with multiple labels, replicate the node on one side (or both) according to the number of edge labels, and assign one label to each resulting edge. Again, SREs connect replicas of the same original node.
The result is a simple graph where all nodes and edges have exactly one label. The SREs ensure that the original structure can be recovered.

*Figure: The process of converting a rich graph (left) into a simple translated graph (right). SREs are shown as dashed lines.*
**Step 3: Mine discriminative subgraphs**
Apply Top-K LEAP on the translated simple graphs. The objective function is information gain – a subgraph is discriminative if its presence/absence strongly separates successful from failed projects.
**Step 4: Reverse translation**
Merge nodes connected by SREs (union of labels), and merge edges that become parallel (union of edge labels). This yields a rich subgraph pattern.
**Soundness & Completeness**: The authors prove that the translation is *sound* (every reverse-translated subgraph is discriminative) and *complete* (every discriminative rich subgraph can be mined). Details are in the technical report [1].
## Experiments and Results
**Dataset:** 64 monthly snapshots of SourceForge.Net (Feb 2005 – May 2010). Projects with >100,000 downloads were labeled *successful*, those with <100 downloads *failed*. After filtering (at least 2 developers, existence at start of study): 224 successful, 3,826 failed projects.
**Classifier:** SVM (LibSVM) with binary features indicating presence/absence of each mined pattern.
**Key Results:**
- **Accuracy: 94.99%** (10-fold cross-validation)
- **AUC: 0.86**
- Translation runs in ~11 seconds total for all projects.
- Graph size grows by factor ~8.4 (theoretical max = 9).
**Top-20 Discriminative Patterns:**
The most informative patterns reflect a clear story:
| # | Pattern (simplified) | Success (%) | Fail (%) | Score |
|---|----------------------|-------------|----------|-------|
| 1 | Two devs, both with 0 PSP and 0 PSC | 26.34 | 92.42 | 0.0952 |
| 8 | Single dev with 1 PSP | 70.98 | 8.05 | 0.0861 |
| 20 | Single dev with 0 PSP | 53.57 | 95.22 | 0.0523 |
*Full table in Figure 10 of the paper.*

**Observations:**
- The #1 pattern (two inexperienced developers with no successful collaborations) appears in 92% of failed projects but only 26% of successful ones – a strong failure indicator.
- Patterns involving a developer with ≥1 past successful project or a pair with ≥1 past successful collaboration are highly predictive of success (occur in 60–70% of successful projects, <6% of failed ones).
- Length of membership (LOM) and length of collaboration history (LCH) do not appear in the top-20 patterns, suggesting they have weaker discriminative power.
## Critical Analysis & Conclusion
**Strengths:**
- The translation framework is theoretically grounded (sound and complete) and practical (linear growth in graph size).
- The mined patterns are interpretable – they directly reveal sociological factors: past success breeds future success; past failure (especially combined with inexperience) is a strong risk signal.
- The approach is fully automatic, requiring no manual feature engineering.
**Limitations:**
- Only tested on SourceForge.Net; external validity to other platforms (e.g., GitHub, enterprises) is not yet shown.
- Success is defined solely by download count – other dimensions (e.g., code quality, user satisfaction) are ignored.
- The study assumes stable developer identities (unique usernames) and does not handle identity merging across projects (though this is a common issue in OSS mining).
**Future Work:**
- Extend to additional success metrics (commits, stars, citations).
- Apply to industrial datasets to validate generalizability.
- Combine the mined patterns with other feature types (e.g., code metrics) for multi-modal prediction.
**Takeaway:** This paper elegantly bridges graph mining and software engineering by showing that *who you worked with before* and *how successful you were* are strong predictors of future project outcomes – a finding that could inform project planning, team formation, and risk assessment tools.
