Beyond Static Exams: Measuring the Quality of Crowdsourced Questions
A Method for Assessing User-generated Tests for Online Courses Exploiting Crowdsourcing Concept
This paper introduces a method for assessing user-generated test questions in online courses by leveraging crowdsourcing. It proposes a recursive scoring mechanism based on "Difficulty" and "Goodness" factors and provides a heuristic iterative algorithm to ensure convergence in question valuation.
TL;DR
To move away from memory-based learning in online courses (MOOCs), researchers at the Kyushu Institute of Technology have developed a system that allows students to generate their own test questions. The core innovation is a mathematical framework that automatically weights these questions based on Difficulty and Goodness, using a recursive algorithm that ensures fair assessment even when question quality varies wildly.
The "Static Pool" Problem
Most online education platforms today suffer from "test fatigue." Teachers create a finite pool of questions that eventually leak online, leading students to memorize answers rather than understand concepts. Crowdsourcing offers a solution by providing a constant stream of new content, but it introduces a major risk: How do we trust a question written by a student? An unfairly difficult or poorly phrased question could ruin a learner's grade.
The Insight: Recursive Quality Control
The authors argue that a question's value is not just about how many people get it wrong. They define two critical metrics:
- Difficulty: Calculated as the complement of the mean score. If everyone gets it right, difficulty is low.
- Goodness: This is the "reputation" of the question, derived from the historical performance (Assessment) of the author.
The technical challenge is that Assessment depends on Goodness, and Goodness depends on Assessment. To solve this circular dependency, the paper proposes a heuristic iterative algorithm.
The Methodology
The system treats assessment as a weighted sum of the raw score, the question's difficulty, and the author's goodness.
The high-level concept of the crowdsourcing feedback loop.
The algorithm starts by assuming an initial score and then cycles through the formulas for Difficulty and Goodness until the values stabilize (converge). This ensures that a single "bad" question won't permanently skew the results, as the iteration balances the weights over time.
Experimental Validation
Using R-based simulations, the researchers tested two main scenarios:
- Variable Goodness: They found that as the algorithm iterates, the "Goodness" factor clearly separates high-quality contributors from low-quality ones.
- Variable Difficulty: They proved that even if questions vary in difficulty, the "Goodness" metric remains a stable indicator of the author's underlying contribution quality.
Simulation results showing the convergence of factors over multiple iterations.
Critical Insight & Future Outlook
The beauty of this approach lies in its Inductive Bias: it assumes that good creators produce good content, and good content is validated by the collective performance of the crowd.
However, there are limitations. The current model assumes scores are normally distributed and doesn't account for "collusion" where groups of users might intentionally fail or pass questions to game the system. Future work will need to move this from simulation to a live e-learning environment like Coursera or edX to test its resilience against real-world user behavior.
Conclusion
This research provides a robust mathematical foundation for decentralized education. By treating test generation as a crowdsourcing task and applying recursive filtering, we can create more dynamic, challenging, and fair learning environments.
