Crowdsourcing Knowledge Components: Can the Crowd Replace the Expert?

Towards Crowdsourcing the Identification of Knowledge Components

2020-07-31
Steven Moore, Huy Anh Nguyen, John C. Stamper
Summary
Problem
Method
Results
Takeaways
Abstract

This paper explores "Towards Crowdsourcing the Identification of Knowledge Components," a study investigating whether non-expert crowdworkers can identify the latent skills (Knowledge Components or KCs) required to solve educational problems. Using the Amazon Mechanical Turk platform across Math and English Writing domains, the authors demonstrate that crowdworkers can successfully match approximately 33% of expert-defined KCs.

TL;DR

Building adaptive educational software requires "Knowledge Component" (KC) mapping—a tedious task usually reserved for experts. This paper investigates if the "crowd" can do it instead. Research shows that while crowdworkers aren't perfect, they can identify about a third of the skills experts do, providing a promising baseline for scaling educational technology without the SME bottleneck.

The Expert Bottleneck in Ed-Tech

In the world of Intelligent Tutoring Systems (ITS), the system needs to know what a student is learning. This is done through KC modeling: if a student fails a geometry problem, the system must know if the failure was in "calculating area" or "understanding subtraction."

Currently, defining these KCs is a manual, labor-intensive process performed by cognitive scientists and subject matter experts. It doesn't scale. The authors of this paper ask a provocative question: Can we use Amazon Mechanical Turk to "map" these skills for us?

The Experiment: Math, Writing, and Priming

The researchers tested two domains:

  1. Mathematics: Middle school level (calculating the area of a wall with windows).
  2. English Writing: Undergraduate level (sentence revision and clause structure).

They also tested Priming. The intuition was that if a crowdworker solved a few related problems first, their "mental muscles" would be warmed up, leading to better KC identification.

Table 1: Expert Generated KCs Table 1: The gold standard KCs created by experts used as a baseline.

Methodology: Coding the Crowd

The team collected 381 responses and used a rigorous qualitative coding process (achieving a high Cohen’s kappa of >0.90) to categorize worker input into three tiers:

  • Irrelevant: Vague or nonsensical.
  • Relevant: Technically needed for the problem, but perhaps not at the right "granularity."
  • Direct Match: Exactly what the expert said.

Key Findings: The Crowd's Performance

The results were surprisingly consistent across domains:

  • The 33% Rule: In both Math and Writing, about 1/3 of worker-generated skills were "Direct Matches" to expert KCs.
  • Math vs. Writing: Workers were better at finding "Relevant" skills in Math (72-78%) than in Writing (54-59%), likely because Math skills are more concrete.
  • The Priming Paradox: Solving "warm-up" problems had no significant effect. In Math, the priming was too easy; in Writing, it was too hard.

Table 5: Direct Match Results Table 5: Distribution of Direct Match KCs across conditions.

Critical Insight: Why did they miss some?

Some expert KCs, like "Id-clause" or "Compose-by-addition," received zero hits from the crowd. Why? Because these terms are "expert-speak." The crowd uses natural language (e.g., "finding the area").

This suggests that while the crowd is great at identifying the functional skills, they struggle with the formal terminology of learning science.

Conclusion & Future Look

This study proves that crowdsourced KC modeling is feasible as a "first draft."

Limitations: The manual analysis of crowd responses currently takes longer than just having an expert do the KCs. The next frontier is automation. If we can use NLP to automatically group and validate these crowdsourced skills, we could build adaptive courses at a fraction of the current cost and time.

The ultimate goal? Learnersourcing—asking students themselves to identify the skills they are using as they learn, creating a self-evolving map of human knowledge.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize "learnersourcing" or student-generated content to build Knowledge Component models in Intelligent Tutoring Systems.
  • Which baseline papers established the "Knowledge-Learning-Instruction" (KLI) framework, and how does this paper's definition of a KC align with that framework?
  • Explore how Natural Language Processing (NLP) or LLMs are currently being used to automate the qualitative coding of crowdsourced educational data.
Contents
Crowdsourcing Knowledge Components: Can the Crowd Replace the Expert?
1. TL;DR
2. The Expert Bottleneck in Ed-Tech
3. The Experiment: Math, Writing, and Priming
4. Methodology: Coding the Crowd
5. Key Findings: The Crowd's Performance
6. Critical Insight: Why did they miss some?
7. Conclusion & Future Look