Unmasking Code Collaboration: Identifying Programmer Social Networks via ARPaD

Identifying Social Networks of Programmers using Text Mining for Code Similarity Detection

2020-12-07
Konstantinos F. Xylogiannopoulos, Panagiotis Karampelas
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces a novel source code similarity detection method that identifies "social networks" of programmers by combining the LERP-RSA data structure and the ARPaD algorithm. It achieved high-efficiency pattern discovery across 46 Python projects, mapping collaborative clusters with O(mn log n) complexity.

TL;DR

Researchers have developed a highly efficient method to detect code similarity and map "social networks" of collaboration between programmers. By leveraging the LERP-RSA data structure and the ARPaD algorithm, the system identifies all repeated patterns in time, proving resilient against common plagiarism "attacks" like variable renaming and code insertion.

Background: The Attribution Problem in a Connected World

With the explosion of platforms like GitHub and StackOverflow, code reuse has become ubiquitous. While beneficial for productivity, it creates a "Source Code Attribution" crisis. Whether it's a student bypassing an assignment's logic or a developer embedding insecure "wild" code into commercial apps, detecting similarity is no longer just about string matching—it's about understanding the algorithmic logic and the social links between authors.

The Core Challenge: Obfuscation and Complexity

Existing methods like n-grams or basic string-matching face two major hurdles:

  1. Sensitivity to Attacks: Students often use "insertion attacks"—adding a single character or changing an operator—to break the "fingerprint" of a code snippet.
  2. Computational Cost: A brute-force one-to-one comparison across a large dataset scales at , making it unusable for big data.

Methodology: From Code Suffixes to Social Graphs

The authors propose a multi-stage pipeline that shifts the focus from simple text to structural patterns.

1. Data Cleansing

To find "logic" rather than "syntax," the authors strip the code of variable names, function names, and formatting. Only reserved words (e.g., if, for, def) and symbols remain. This normalizes the code, making "Variable X" and "Variable Y" identical in the eyes of the algorithm.

2. The LERP-RSA and ARPaD Engine

At the heart of the system is the Longest Expected Repeated Pattern Reduced Suffix Array (LERP-RSA).

  • How it works: It treats all code snippets as part of a massive multivariate data structure.
  • The ARPaD Advantage: Unlike n-gram approaches that require a fixed "n" (length), the All Repeated Patterns Detection (ARPaD) algorithm finds all patterns between the Shorter Pattern Length (SPL) and the Longest Expected Repeated Pattern (LERP) in a single pass.

Model Architecture and Data Flow Fig 1: Visual representation of common patterns (colored bars) overlaid on code sequences.

Experimental Insights: Mapping the Network

The researchers tested their method on a new dataset of 46 Python assignments. By setting different thresholds for "overlapping length," they could filter out noise and reveal clear "islands" of collaboration.

Breaking the Insertion Attack

A standout result was the comparison of sequences 40 and 45. While a single character was inserted to break the code, ARPaD detected two patterns (87 and 48 characters long). Combined, they achieved nearly 100% coverage, proving a direct copy-paste relationship that simpler tools might miss.

Social Network Graphs Fig 2: Social network graphs visualizing collaboration clusters at a total overlapping length threshold of ≥ 60.

Critical Analysis & Conclusion

Takeaways

  • Efficiency: The complexity allows this to run on a standard commodity laptop, making it accessible for classroom or small-firm use.
  • Graph-Based Insights: By visualizing code similarity as a social network, investigators can identify "influencers" (source code authors) and "followers" (those who copy).

Limitations & Future Work

The current method relies heavily on reserved words. While effective for Python, its performance might vary in languages with fewer reserved keywords or highly boilerplate-heavy frameworks. The authors suggest that future versions will include percentage-based weights to create directed graphs—indicating not just if a connection exists, but who likely copied from whom based on original sequence length.

This research bridges the gap between text mining and social network analysis, providing a robust framework for preserving integrity in the open-source era.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Suffix Arrays or Suffix Trees for cross-platform binary code similarity detection beyond academic plagiarism.
  • Which original study first proposed the LERP-RSA data structure, and how does its space complexity compare to traditional Suffix Arrays in big data contexts?
  • Explore research that applies Social Network Analysis (SNA) metrics like centrality or community detection to developer collaboration patterns in large-scale GitHub repositories.
Contents
Unmasking Code Collaboration: Identifying Programmer Social Networks via ARPaD
1. TL;DR
2. Background: The Attribution Problem in a Connected World
3. The Core Challenge: Obfuscation and Complexity
4. Methodology: From Code Suffixes to Social Graphs
4.1. 1. Data Cleansing
4.2. 2. The LERP-RSA and ARPaD Engine
5. Experimental Insights: Mapping the Network
5.1. Breaking the Insertion Attack
6. Critical Analysis & Conclusion
6.1. Takeaways
6.2. Limitations & Future Work