Framework for Community Modeling: Bridging Crowdsourcing Architecture and User Behavior
Towards a framework for supporting community modeling in crowdsourcing systems
The paper proposes a high-level framework to support community modeling in Crowdsourcing Systems (CS). It establishes a systematic mapping between four types of crowdsourcing (Processing, Rating, Creation, Solving) and eight classes of User-Model Classes (UMCs) to assist developers in choosing optimal user modeling strategies.
TL;DR
Understanding the "crowd" in crowdsourcing is often treated as a monolith, leading to poor retention and recruitment. This paper introduces a high-level framework that categorizes Crowdsourcing Systems (CS) and maps them to specific User-Model Classes (UMC). By aligning system goals with user modeling dimensions like Implicit vs. Explicit data collection, the framework provides a roadmap for researchers to design more adaptive and motivating systems.
Problem & Motivation: The Diversity Paradox
The strength of crowdsourcing lies in its diversity, but this same diversity makes it nearly impossible to design a single interface or incentive structure that works for everyone. The authors identify a gap: researchers often lack the theoretical tools to decide how to model their users. Should we track every individual's behavior (Individual Model) or focus on the average user (Canonical Model)? Should we ask users for data directly (Explicit) or watch from the background (Implicit)? This paper argues that the answer depends entirely on the type of crowdsourcing system being built.
Methodology: The Two-Stage Selection Logic
The proposed framework operates in two distinct logical steps, bridging the gap between System Theory and Human-Computer Interaction (HCI).
1. Pre-selection (Classification)
Using a flowchart, the framework forces developers to answer key questions: Who is doing the task? Why? and How? This classifies the system into one of four prototypes:
- Crowd Processing: High quantity, low intellectual effort (e.g., data labeling).
- Crowd Rating: Collective evaluation for accuracy (e.g., product reviews).
- Crowd Creation: Focus on singularity and individual impact (e.g., logo design).
- Crowd Solving: Solving specific, isolated problems (e.g., bug hunting).

2. Matching UMCs to System Types
Once the system is classified, the framework suggests relevant User-Model Classes (UMCs). For example, in Crowd Solving, the individual matters most, so Individual Models are recommended. In Crowd Rating, where the "wisdom of the crowd" is an aggregate, Canonical Models suffice.
| Dimension | Meaning in Crowdsourcing Context |
|---|---|
| Canonical vs. Individual | Does the system treat the user as a "type" or a unique entity? |
| Explicit vs. Implicit | Does the system require user input (high friction) or use sensor/log data (low friction)? |
| Long-term vs. Short-term | Are we modeling permanent traits (preferences) or temporary states (current location)? |

Experiments & Results: Real-World Validation
The authors validated the framework using two real-world projects:
- Demorô: A participatory sensing system for monitoring service waiting times.
- Acropolis: A social curation platform.
The study found that the leaders of these projects were able to quickly identify that Implicit, Canonical, Long-term models were the most viable for their business models. Why? Because expecting a "crowd" to explicitly update profiles (Explicit) creates too much friction, and tracking every micro-move (Individual) in a high-volume sensing app is often computationally unnecessary compared to a canonical archetype.
Critical Analysis & Conclusion
Takeaway
The framework moves crowdsourcing research from "intuition-based" design to "strategy-based" design. It clarifies the trade-offs: choosing an Individual-Explicit model offers high precision but risks exhausting the user, while Canonical-Implicit models offer scale at the cost of personalization.
Limitations & Future Work
The preliminary study noted that the flowchart in the pre-selection stage was somewhat difficult for researchers to navigate (scoring 3-4 on the difficulty scale). Future versions of the framework need to simplify technical jargon. Furthermore, the goal is to drive the framework down to a "lower level," providing specific algorithms (e.g., Bayesian networks or personas) for each UMC.
Ultimately, this work acts as a foundational "decision support system" for anyone looking to build more human-centric crowdsourcing platforms.
