Beyond the Answer: Modeling Knowledge Through Learnersourcing Contributions

Modelling Learners in Crowdsourcing Educational Systems

2020-01-01
Solmaz Abdi, Hassan Khosravi, Shazia Sadiq
Summary
Problem
Method
Results
Takeaways
Abstract

The paper proposes an encoding extension for Knowledge Tracing Machines (KTMs) to model students' knowledge states by integrating "learnersourcing" activities. Beyond traditional assessment performance, the model incorporates data from student contributions like content creation and moderation, achieving superior predictive accuracy (AUC improvement) on datasets from the RiPPLE platform.

TL;DR

Traditionally, AI-driven educational systems judge a student's knowledge by whether they get an answer "right" or "wrong." This paper challenges that narrow view by introducing a multi-task Knowledge Tracing Machine (KTM). By accounting for "learnersourcing"—the act of students creating and moderating content—the model significantly improves its ability to predict a student's true competency.

Positioning: This work moves beyond traditional "Knowledge Tracing" into the realm of Active Learning Analytics, proving that how much you contribute is as telling as how much you correctly answer.

The "Assessment-Only" Bottleneck

In traditional Intelligent Tutoring Systems (ITS), the data loop is simple:

  1. System presents a question.
  2. Student answers.
  3. Model (like BKT or IRT) updates the knowledge state.

However, modern crowdsourcing platforms like RiPPLE allow students to act as teachers: they create questions, moderate submissions, and leave feedback. Prior work is often blind to these "higher-order" learning activities. The authors argue that a student who can create a quality question about "Biological Fate of Drugs" likely possesses a deeper mastery than one who simply guesses the right answer on a multiple-choice item.

Methodology: Extending Knowledge Tracing Machines

The authors utilize the Knowledge Tracing Machine (KTM) framework, which uses Factorization Machines to model interactions between students, items, and concepts.

The Multi-Task Extension

The core innovation is the re-encoding of the input vector. Instead of a single "opportunity" count for attempts, they decompose student activity into specific task types .

  • : Opportunities on Attempting items.
  • : Opportunities on Creating items.
  • : Opportunities on Moderating items.

This allows the model to assign different weights to different types of engagement. Creating a resource is mathematically treated as a different "signal" of knowledge than merely attempting one.

Model Encoding and Multi-task Architecture Figure 1: Comparison of traditional KTM input vs. the proposed multi-task encoding (Encoding interactions chronologically across Attempt, Create, and Moderate tasks).

Experiments and Results

The model was tested on two real-world datasets: Medi (Medical Licensing) and Pharm (Pharmacy).

Key Findings:

  1. Superior Accuracy: The proposed model (m5) achieved the highest AUC (0.723 in Medi) compared to IRT, AFM, and PFA.
  2. The "Creation" Signal: Including (Create) and (Moderate) consistently bettered the baseline that only looked at attempts ().
  3. Validation of Bloom’s Taxonomy: The results provide empirical evidence for learning theories suggesting that higher-order tasks (creating/evaluating) are more indicative of deep understanding than lower-order tasks (recalling/applying).

Experimental Results Comparison Table 2: Performance comparison showing m5 (multi-task) outperforming traditional models across Accuracy and AUC.

Critical Insight & Future Outlook

Why it matters

This paper provides a blueprint for the next generation of Open Learner Models (OLMs). If a student sees that their competency score increases when they help moderate a peer's question, they are more likely to engage in that community-driven behavior. It solves the "cold start" problem of student participation by quantifying the pedagogical value of their contributions.

Limitations

  • Dataset Scale: The study used relatively small datasets (131–179 students). Whether these specific weights for "creation" generalize across less technical disciplines (e.g., Humanities vs. Medicine) remains to be seen.
  • Quality vs. Quantity: Currently, the model tracks the number of opportunities. A future iteration should incorporate the quality of the created content (e.g., was the created question actually good?) to further refine the knowledge estimate.

Conclusion

By integrating learnersourcing into the mathematical framework of Knowledge Tracing, we can finally treat the student not just as a consumer of information, but as an active participant whose contributions are a vital mirror of their expertise.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Factorization Machines or Knowledge Tracing Machines (KTM) for student performance prediction in collaborative learning environments.
  • What are the seminal works on the pedagogy of "learnersourcing," and how do they differentiate between the learning gains of content creation versus content consumption?
  • Explore research that integrates multi-modal interaction data (e.g., forum posts, video annotations) into Knowledge Tracing models to improve adaptivity in MOOCs.
Contents
Beyond the Answer: Modeling Knowledge Through Learnersourcing Contributions
1. TL;DR
2. The "Assessment-Only" Bottleneck
3. Methodology: Extending Knowledge Tracing Machines
3.1. The Multi-Task Extension
4. Experiments and Results
4.1. Key Findings:
5. Critical Insight & Future Outlook
5.1. Why it matters
5.2. Limitations
5.3. Conclusion