Unmasking Code Collaboration: Identifying Programmer Social Networks via ARPaD
Identifying Social Networks of Programmers using Text Mining for Code Similarity Detection
The paper introduces a novel source code similarity detection method that identifies "social networks" of programmers by combining the LERP-RSA data structure and the ARPaD algorithm. It achieved high-efficiency pattern discovery across 46 Python projects, mapping collaborative clusters with O(mn log n) complexity.
TL;DR
Researchers have developed a highly efficient method to detect code similarity and map "social networks" of collaboration between programmers. By leveraging the LERP-RSA data structure and the ARPaD algorithm, the system identifies all repeated patterns in time, proving resilient against common plagiarism "attacks" like variable renaming and code insertion.
Background: The Attribution Problem in a Connected World
With the explosion of platforms like GitHub and StackOverflow, code reuse has become ubiquitous. While beneficial for productivity, it creates a "Source Code Attribution" crisis. Whether it's a student bypassing an assignment's logic or a developer embedding insecure "wild" code into commercial apps, detecting similarity is no longer just about string matching—it's about understanding the algorithmic logic and the social links between authors.
The Core Challenge: Obfuscation and Complexity
Existing methods like n-grams or basic string-matching face two major hurdles:
- Sensitivity to Attacks: Students often use "insertion attacks"—adding a single character or changing an operator—to break the "fingerprint" of a code snippet.
- Computational Cost: A brute-force one-to-one comparison across a large dataset scales at , making it unusable for big data.
Methodology: From Code Suffixes to Social Graphs
The authors propose a multi-stage pipeline that shifts the focus from simple text to structural patterns.
1. Data Cleansing
To find "logic" rather than "syntax," the authors strip the code of variable names, function names, and formatting. Only reserved words (e.g., if, for, def) and symbols remain. This normalizes the code, making "Variable X" and "Variable Y" identical in the eyes of the algorithm.
2. The LERP-RSA and ARPaD Engine
At the heart of the system is the Longest Expected Repeated Pattern Reduced Suffix Array (LERP-RSA).
- How it works: It treats all code snippets as part of a massive multivariate data structure.
- The ARPaD Advantage: Unlike n-gram approaches that require a fixed "n" (length), the All Repeated Patterns Detection (ARPaD) algorithm finds all patterns between the Shorter Pattern Length (SPL) and the Longest Expected Repeated Pattern (LERP) in a single pass.
Fig 1: Visual representation of common patterns (colored bars) overlaid on code sequences.
Experimental Insights: Mapping the Network
The researchers tested their method on a new dataset of 46 Python assignments. By setting different thresholds for "overlapping length," they could filter out noise and reveal clear "islands" of collaboration.
Breaking the Insertion Attack
A standout result was the comparison of sequences 40 and 45. While a single character was inserted to break the code, ARPaD detected two patterns (87 and 48 characters long). Combined, they achieved nearly 100% coverage, proving a direct copy-paste relationship that simpler tools might miss.
Fig 2: Social network graphs visualizing collaboration clusters at a total overlapping length threshold of ≥ 60.
Critical Analysis & Conclusion
Takeaways
- Efficiency: The complexity allows this to run on a standard commodity laptop, making it accessible for classroom or small-firm use.
- Graph-Based Insights: By visualizing code similarity as a social network, investigators can identify "influencers" (source code authors) and "followers" (those who copy).
Limitations & Future Work
The current method relies heavily on reserved words. While effective for Python, its performance might vary in languages with fewer reserved keywords or highly boilerplate-heavy frameworks. The authors suggest that future versions will include percentage-based weights to create directed graphs—indicating not just if a connection exists, but who likely copied from whom based on original sequence length.
This research bridges the gap between text mining and social network analysis, providing a robust framework for preserving integrity in the open-source era.
