Crowdsourcing the Review: Can a Global Crowd Solve the Research Bottleneck in SE?
Crowdsourcing in Systematic Reviews: A Systematic Mapping and Survey
This study explores the application of crowdsourcing to Systematic Reviews (SRs) in Software Engineering (SE) through a Systematic Mapping (SM) and a survey of 39 researchers. It identifies core benefits such as reduced conduction time and bias, while proposing a four-phase crowd-based process (Training, Selection, Activity, Aggregation) to handle labor-intensive SR tasks like paper screening and data extraction.
Executive Summary
TL;DR: This paper tackles the "labor crisis" in Evidence-Based Software Engineering by analyzing how crowdsourcing can transform the Systematic Review (SR) process. Through a systematic mapping and researcher survey, the authors propose a 4-phase framework to shift repetitive screening and extraction tasks from highly-paid researchers to a distributed "crowd," potentially cutting down review times significantly while maintaining accuracy through strict quality-control loops.
Positioning: This is a foundational survey and methodology-shaping work that bridges the gap between Citizen Science paradigms and the rigid requirements of Evidence-Based Software Engineering.
Motivation: The Scale Problem in Science
Modern Software Engineering moves at a breakneck speed, and the volume of primary studies is growing exponentially. Traditional SRs are becoming a bottleneck—a single review can consume over 120 hours of expert time. The authors argue that we are wasting "expert brainpower" on binary classification tasks (Include/Exclude) that could be handled by a managed crowd and verified through consensus.
Methodology: The 4-Phase Crowd-Review Engine
To address concerns over quality and reliability, the authors synthesized a structured process designed to filter out noise and incompetence.
1. The Workflow Architecture
The proposed process doesn't just "throw tasks over the wall." It implements a systematic pipeline:
- Phase 1: Training: Workers are presented with "Gold Standard" examples (clear-cut inclusion and exclusion cases).
- Phase 2: Selection (The Hurdle): Workers must pass a "qualification test" with >70% accuracy on known papers before being allowed to touch real data.
- Phase 3: Crowd-Activity: Monitoring is continuous. "Trap questions" are hidden in the workflow to ensure workers stay attentive.
- Phase 4: Aggregation: Final decisions are reached via majority voting, reducing the impact of individual human error.
Figure 1: The standard four-phase cycle for crowd-based evidence synthesis.
Experiments and Researcher Sentiment
The study’s survey of 39 active SE researchers provided a reality check on this vision:
- The Optimism: Speed was cited as the #1 benefit (76.9%). Interestingly, many researchers believe a broad crowd is less biased than a small, tight-knit group of domain experts who might have "intellectual blinders."
- The Skepticism: Quality control is the "Elephant in the room." 76.9% of experts worried about task accuracy, and 56.4% were concerned that the overhead of managing a crowd might outweigh the time saved.
Figure 2: The study utilized a rigorous three-stage search (Automatic, Manual, and Snowballing) to map the current state of knowledge.
Critical Analysis & Conclusion
The Takeaway
Crowdsourcing in SE Systematic Reviews is effectively in its "Alpha" stage. While the theoretical speed gains are massive, the practical infrastructure (specialized platforms for SR) is currently missing. Most researchers are using general-purpose tools like Amazon Mechanical Turk, which aren't built for the nuances of academic literature.
Limitations & Future Work
- The "Expertise" Paradox: While simple screening can be outsourced, high-level synthesis still requires years of training. The "Shortest Run" algorithm mentioned in the paper—where AI predicts when a paper is too complex for the crowd and routes it back to the expert—is the most promising path forward.
- Incentives: The paper notes a conflict in motivation. While some look for financial gain, many in the academic community prefer "social credit" (co-authorship or acknowledgment) as a reward.
Final Thought: If we want "Evidence-Based" SE to survive the data deluge, we must stop treating SRs as a "monastic" individual task and start treating them as a collaborative, distributed data-engineering problem.
