The Crowd as a Compiler: Redefining Software Engineering Through Crowdsourcing

Crowdsourcing in software engineering: models, motivations, and challenges

2019-05-27
D. Thomas
Summary
Problem
Method
Results
Takeaways

This article surveys the landscape of Crowdsourcing in Software Engineering, categorizing it into three primary models: Peer Production, Competitions, and Microtasking. It establishes an eight-dimensional framework to evaluate these models and discusses how they provide scalability and diverse solutions for modern software development.

TL;DR

Crowdsourcing is no longer just for protein folding or Wikipedia; it is fundamentally disrupting software engineering. By leveraging an open call to an "undefined" network of global talent, organizations are achieving massive parallelism in bug hunting, UI design, and development. This article explores the core models of this shift—Peer Production, Competitions, and Microtasking—and identifies the 8 dimensions that define the success of a crowdsourced project.

The Shift from In-House to Out-of-Bounds

Historically, software was built behind closed doors. The primary limitation was always bandwidth—a fixed number of developers can only write a fixed amount of code. Previous outsourcing models tried to solve this but were bogged down by rigid contracts and high management overhead.

The authors argue that the "Crowd" offers a different physical intuition: Collective Intelligence. Instead of a pipeline, think of a massive, elastic cloud of specialists who engage on-demand. The motivation isn't just money (extrinsic); it's often reputation and learning (intrinsic), which changes the quality and diversity of the output.

Mapping the Landscape: The 8 Dimensions

To understand how to apply crowdsourcing, the authors introduce a rigorous taxonomy. It isn't enough to say "we are crowdsourcing"; one must define where the project sits across these axes:

  • Task Interdependence & Context: These are the "killers" of crowdsourcing. If a dev needs to understand 1 million lines of code to fix one bug, the crowd fails.
  • Locus of Control: Is the client dictating the micro-steps (TopCoder), or is the worker choosing the direction (Open Source)?

The 8 Dimensions of Crowdsourcing

Three Archetypes of Crowdsourced Work

1. Peer Production (The Open Source Way)

  • Core Logic: High expertise, high context, but decentralized control.
  • Advantage: Long-term sustainability and high "intrinsic" motivation (e.g., Linux, Rails).

2. Competitions (The TopCoder Model)

  • Core Logic: Parallelism through rivalry.
  • Advantage: The client receives multiple solutions to the same problem and only pays for the winner. This leads to higher quality through "Natural Selection" of code architectures.

3. Microtasking (The Atomic Unit)

  • Core Logic: Breaking work into 5-minute chunks.
  • Advantage: Extreme scalability. Using platforms like UserTesting.com, a company can fix a vulnerability or test a browser on 100 systems in just two hours.

Comparison of Crowdsourcing Models

The "Decomposition" Challenge: Why Crowdsourcing is Hard

The most profound insight in this paper is the Decomposition Hurdle. Unlike tagging images, software development is highly interdependent.

The authors suggest that the future of Software Engineering (SE) isn't just better coding, but better orchestration. We need workflows that can:

  1. Atomize Complexity: Turn a feature request into 50 independent microtasks.
  2. Ensure Quality: Since the workers are "unknown," the system must use "Replication" (having multiple people do the same task) or "Verification Games" to ensure the code works.

Final Analysis: Future or Fantasy?

Is the future of SE a world of "mindless tokens" doing micro-tasks, or "highly skilled freelancers" pursuing passion? The reality will likely be a hybrid.

The paper concludes that while crowdsourcing has mastered the "peripheral" (testing, Q&A like StackOverflow), the final frontier is the "core"—building entire enterprise systems through a crowd. To get there, we don't just need more developers; we need a new "Operating System" for how work is distributed across the planet.

Takeaways for the Industry:

  • For Managers: Start with "High Replication/Low Context" tasks like compatibility testing and bug bounties.
  • For Researchers: Focus on automated task decomposition and context-reduction techniques.

Find Similar Papers

Try Our Examples

  • Search for recent empirical studies or systematic literature reviews on the effectiveness of microtask programming in large-scale software projects.
  • Which seminal papers first defined the "decomposition problem" in crowdsourced software engineering, and what algorithmic approaches have been proposed to automate task splitting?
  • Explore how Large Language Models (LLMs) are currently being integrated into crowdsourcing platforms to reduce the "task context" burden for human workers.
Contents
The Crowd as a Compiler: Redefining Software Engineering Through Crowdsourcing
1. TL;DR
2. The Shift from In-House to Out-of-Bounds
3. Mapping the Landscape: The 8 Dimensions
4. Three Archetypes of Crowdsourced Work
4.1. 1. Peer Production (The Open Source Way)
4.2. 2. Competitions (The TopCoder Model)
4.3. 3. Microtasking (The Atomic Unit)
5. The "Decomposition" Challenge: Why Crowdsourcing is Hard
6. Final Analysis: Future or Fantasy?
6.1. Takeaways for the Industry: