KT4Crowd: Predicting Expert Performance by Tracing Worker Knowledge Evolution
Predicting Crowdsourcing Worker Performance with Knowledge Tracing
This paper introduces KT4Crowd, a novel framework that adapts Knowledge Tracing (KT) models—originally designed for Intelligent Tutoring Systems—to predict worker performance in Knowledge-Intensive Crowdsourcing (CKI-C). By utilizing the DKVMN model within this framework, the authors achieve State-of-the-Art (SOTA) results in predicting developer success on platforms like Topcoder.
TL;DR
Predicting who will win a high-stakes software competition is notoriously difficult because expert skills are not static. This paper introduces KT4Crowd, a framework that treats competitive crowdsourcing like an Intelligent Tutoring System. By breaking down complex tasks and applying Knowledge Tracing (KT) models (specifically DKVMN), the authors can track how a worker's "knowledge state" evolves over time, leading to performance predictions that significantly outperform traditional Elo-style rating systems.
Background: The Problem with Static Ratings
In competitive platforms like Topcoder or Kaggle, workers aren't just labels; they are learners. Traditional worker models often assume a fixed "ability score" (like a chess rating). However, in Knowledge-Intensive Crowdsourcing (KI-C):
- Skills Evolve: A developer's proficiency in "Java" or "Machine Learning" changes with every task they complete.
- Multi-Dimensional Tasks: A single task might require UI design, algorithm optimization, and database management simultaneously.
- No Standard Answer: Unlike a math quiz, performance is measured in relative ranks and continuous scores.
The Insight: Crowdsourcing as a "Classroom"
The authors realized that CKI-C mirrors Intelligent Tutoring Systems (ITS). In both cases, we want to know: If a user has done X and Y in the past, can they solve Z now? To bridge the gap between educational KT and complex crowdsourcing, the KT4Crowd framework introduces two critical adaptations.
Methodology: The KT4Crowd Framework
The core innovation lies in how data is "translated" for the AI:
- Task Decomposition: Since KT models (like DKT or DKVMN) usually handle one concept at a time, KT4Crowd breaks a task with features into separate subtasks.
- Results Transformation: Instead of raw scores, the model focuses on "Good Performance"—defined as achieving a rank and score above specific thresholds.
Architecture of the Prediction Process
The framework utilizes a trained KT model to project a worker's future performance. It simulates a "loss" (0) for each subtask, passes it through a Memory Network, and uses the resulting state to predict the probability of success.
Figure 1: The prediction workflow showing the transformation of historical sequences into subtask predictions followed by Majority Voting.
Experiments and Results
The authors tested their framework against traditional rating systems (Glicko-2) and standard KT applications on a massive dataset of 50,625 submissions from Topcoder.
Sequence Prediction Performance
The comparison involved DKT (Deep Knowledge Tracing) and DKVMN (Dynamic Key-Value Memory Networks).
| Metric | DKT-S (Ours) | DKVMN-S (Ours) | DKT-O (Baseline) |
|---|---|---|---|
| AUC | 0.7467 | 0.8412 | 0.7166 |
| Accuracy | 0.7125 | 0.8254 | 0.6980 |
| F1 Score | 0.4807 | 0.7540 | 0.5498 |
Key Takeaways from the Data:
- DKVMN-S is the clear winner: Utilizing a memory-augmented neural network allows the model to store specific skill proficiencies more effectively than a standard LSTM (DKT).
- The Framework Matters: The "-S" (KT4Crowd) versions consistently outperformed the "-O" (One-feature) versions, proving that decomposing tasks into multiple skills is essential for accuracy.
Winner Prediction: Beating the Industry Standard
When it came to predicting specifically who would "win" or perform excellently on a new task, KT4Crowd (with DKVMN) outperformed the actual Topcoder Rating System and the Glicko-2 system by a wide margin (AUC 0.78 vs traditional methods failing to capture the complexity).
Table: Comparison of KT4Crowd against Glicko and Topcoder rating systems.
Critical Insight & Conclusion
Why does this work? Traditional rating systems are summaries; Knowledge Tracing is a history. By maintaining a memory of how a worker handled specific skills in the past, the model can navigate the "Expertise Cold Start" problem and the "Skill Decay" problem simultaneously.
Limitations: The model currently relies on "Majority Voting" for subtasks, which assumes all skills in a task are equally important. Future iterations might benefit from an Attention Mechanism to weight the "Key Feature" of a task more heavily than secondary requirements.
Ultimately, KT4Crowd proves that the boundary between "learning" and "working" is porous. If we can trace how knowledge is acquired, we can predict how it will be applied in the global digital economy.
