Beyond the Answer: Modeling Knowledge Through Learnersourcing Contributions
Modelling Learners in Crowdsourcing Educational Systems
The paper proposes an encoding extension for Knowledge Tracing Machines (KTMs) to model students' knowledge states by integrating "learnersourcing" activities. Beyond traditional assessment performance, the model incorporates data from student contributions like content creation and moderation, achieving superior predictive accuracy (AUC improvement) on datasets from the RiPPLE platform.
TL;DR
Traditionally, AI-driven educational systems judge a student's knowledge by whether they get an answer "right" or "wrong." This paper challenges that narrow view by introducing a multi-task Knowledge Tracing Machine (KTM). By accounting for "learnersourcing"—the act of students creating and moderating content—the model significantly improves its ability to predict a student's true competency.
Positioning: This work moves beyond traditional "Knowledge Tracing" into the realm of Active Learning Analytics, proving that how much you contribute is as telling as how much you correctly answer.
The "Assessment-Only" Bottleneck
In traditional Intelligent Tutoring Systems (ITS), the data loop is simple:
- System presents a question.
- Student answers.
- Model (like BKT or IRT) updates the knowledge state.
However, modern crowdsourcing platforms like RiPPLE allow students to act as teachers: they create questions, moderate submissions, and leave feedback. Prior work is often blind to these "higher-order" learning activities. The authors argue that a student who can create a quality question about "Biological Fate of Drugs" likely possesses a deeper mastery than one who simply guesses the right answer on a multiple-choice item.
Methodology: Extending Knowledge Tracing Machines
The authors utilize the Knowledge Tracing Machine (KTM) framework, which uses Factorization Machines to model interactions between students, items, and concepts.
The Multi-Task Extension
The core innovation is the re-encoding of the input vector. Instead of a single "opportunity" count for attempts, they decompose student activity into specific task types .
- : Opportunities on Attempting items.
- : Opportunities on Creating items.
- : Opportunities on Moderating items.
This allows the model to assign different weights to different types of engagement. Creating a resource is mathematically treated as a different "signal" of knowledge than merely attempting one.
Figure 1: Comparison of traditional KTM input vs. the proposed multi-task encoding (Encoding interactions chronologically across Attempt, Create, and Moderate tasks).
Experiments and Results
The model was tested on two real-world datasets: Medi (Medical Licensing) and Pharm (Pharmacy).
Key Findings:
- Superior Accuracy: The proposed model (m5) achieved the highest AUC (0.723 in Medi) compared to IRT, AFM, and PFA.
- The "Creation" Signal: Including (Create) and (Moderate) consistently bettered the baseline that only looked at attempts ().
- Validation of Bloom’s Taxonomy: The results provide empirical evidence for learning theories suggesting that higher-order tasks (creating/evaluating) are more indicative of deep understanding than lower-order tasks (recalling/applying).
Table 2: Performance comparison showing m5 (multi-task) outperforming traditional models across Accuracy and AUC.
Critical Insight & Future Outlook
Why it matters
This paper provides a blueprint for the next generation of Open Learner Models (OLMs). If a student sees that their competency score increases when they help moderate a peer's question, they are more likely to engage in that community-driven behavior. It solves the "cold start" problem of student participation by quantifying the pedagogical value of their contributions.
Limitations
- Dataset Scale: The study used relatively small datasets (131–179 students). Whether these specific weights for "creation" generalize across less technical disciplines (e.g., Humanities vs. Medicine) remains to be seen.
- Quality vs. Quantity: Currently, the model tracks the number of opportunities. A future iteration should incorporate the quality of the created content (e.g., was the created question actually good?) to further refine the knowledge estimate.
Conclusion
By integrating learnersourcing into the mathematical framework of Knowledge Tracing, we can finally treat the student not just as a consumer of information, but as an active participant whose contributions are a vital mirror of their expertise.
