ICS Caption Editor: Leveraging Crowdsourcing to Bridge the ASR Accuracy Gap in STEM Education
A crowdsourcing caption editor for educational videos
The paper presents the ICS Caption Editor, a web-based crowdsourcing framework designed to generate and refine captions for STEM educational videos. By integrating Automatic Speech Recognition (ASR) with a collaborative student-driven editing workflow, the system achieves near-perfect caption accuracy (99%) for complex technical lectures.
TL;DR
The ICS Caption Editor is a collaborative platform that turns the tedious task of lecture captioning into a manageable crowdsourced activity. By combining imperfect Automatic Speech Recognition (ASR) with student-led micro-tasks, the system achieves 99% accuracy, making technical STEM videos searchable and highly accessible without the high cost of professional transcription.
The "Broken" State of Automated Transcription
While LLMs and modern AI have improved speech-to-text, this paper highlights a persistent struggle in academia: Technical Context. In live STEM lectures, the combination of complex terminology (e.g., "Ungrammatical constructs" due to math variables), classroom echoes, and diverse instructor accents creates a "perfect storm" of errors.
The authors' pre-study found that even commercial tools like Dragon Naturally Speaking and YouTube ASR averaged only 68% to 85% accuracy for live lectures. This makes the resulting captions more of a distraction than a help. Manual professional services are the gold standard but are far too expensive for every-day classroom use.
Methodology: The Power of the Crowd
The core insight of the ICS Caption Editor is that students are the best domain experts for their own courses. If 10-12 students spend 45 minutes each, an 80-minute lecture is captioned perfectly in a few days.
Key Architectural Features:
- Micro-Partitioning: The system breaks the transcript into segments of 5 sentences. This prevents "task fatigue" and allows multiple students to work simultaneously.
- Audio Looping & Variable Speed: To catch difficult phrases, the editor loops the specific audio segment and provides a "PlaySpeed" tool to slow down playback without pitch distortion.
- The "Review" Protocol: Users can mark segments as "Needs Review," creating a multi-pass verification system that ensures technical terms are triple-checked.
Figure 1: The ICS Caption Editor interface featuring synchronized video, text editing, and status tracking (Complete vs. Needs Review).
Experimental Results: Accuracy Meets Efficiency
The evaluation focused on two Computer Science courses. The results were striking:
- Near-Perfect Accuracy: Final captions reached 99% accuracy, far surpassing any ASR-only approach.
- Distributed Effort: While one person would take ~10 hours to caption a lecture, the crowdsourced group completed it with a median individual effort of just 45 minutes.
- Subjective Value: An overwhelming majority of students (most of whom were non-native English speakers) reported that captions improved their Efficiency, Note-taking, and Learning (see Figure 2).
Figure 2: Survey data showing that students perceived significant improvements in learning and attention due to the presence of captions.
Academic Insight: Why it Works
The success of the ICS Editor is rooted in collaborative validation. In technical domains, "Tool Weakness" (incorrect ASR hypothesis) accounts for 50% of errors. By providing students with a visual frame of the lecture slides alongside the audio, they use visual cues to correct what the ASR "heard" incorrectly.
Critical Analysis & Conclusion
Takeaway
The ICS Caption Editor proves that captioning is not just an accessibility requirement but a pedagogical tool. By involving students in the process, the workload is socialized, and the final output becomes a searchable "videobook" that functions like a textbook.
Limitations & Future Work
The study was based on a smaller sample size (24 students). Additionally, the incentive structure—how to keep students motivated to participate consistently—remains an open question, though many indicated a willingness to work for academic credit. Future iterations could benefit from Large Language Models (LLMs) to provide "context-aware" auto-corrections before the students even begin their review, further reducing the manual workload.
Source Context: This work was supported by the National Science Foundation (NSF) and integrated into the ICS Videos framework used at the University of Houston.
