CrowdMOS: Democratizing Subjective Audio Quality Assessment

CROWDMOS: An approach for crowdsourcing mean opinion score studies

2011-05-01
Flavio P. Ribeiro, Dinei A. F. Florêncio, Cha Zhang, Michael L. Seltzer
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces crowdMOS, a cost-effective crowdsourcing framework for conducting Mean Opinion Score (MOS) subjective audio quality studies via Amazon Mechanical Turk. By implementing automated screening and a two-way random effects model, it achieves high-quality results comparable to laboratory studies at a fraction of the cost.

TL;DR

CrowdMOS is an open-source framework that moves Mean Opinion Score (MOS) testing from expensive labs to the internet crowd via Amazon Mechanical Turk. By using a clever Two-Way Random Effects Model to filter out bad actors and calculate honest confidence intervals, the authors prove that $11/algorithm can yield results nearly identical to professional lab studies.

The Motivation: Lab Guilt and Objective Failure

In the world of signal processing, the Mean Opinion Score (MOS) is the gold standard. However, doing it "by the book" (ITU-T P.800) is a nightmare of logistics: soundproof booths, standardized headphones, and pre-screened listeners.

Because of this, researchers often resort to Objective Measures (like PESQ). The problem? These algorithms are often "blind" to modern artifacts like jitter buffer adjustments or packet loss concealment. We need humans, but we need them to be affordable.

The Solution: CrowdMOS

The authors propose crowdMOS. It’s not just "asking people on the internet"; it’s a systematic approach to making crowdsourcing scientifically rigorous.

1. The Strategy for Human Intelligence Tasks (HITs)

To keep workers engaged and honest, the authors designed a specific UI/UX and incentive structure:

  • Incentives over Barriers: Instead of a "qualification test" that scares people away, they use a bonus system based on performance and volume.
  • Throughput Design: Using radio buttons and optimized layouts to ensure a worker can finish a 10-sample task in ~90 seconds.

HIT Throughput

2. The Math: Modeling the Noise

The core technical contribution is treating worker "noise" not as a nuisance, but as a statistical parameter. They use a two-way random effects model:

Where:

  • : Variation in the difficulty of the sentence.
  • : The specific bias/preference of the worker.
  • : Pure subjective uncertainty.

By isolating these variables, they can calculate 95% Confidence Intervals (CIs) that are far more accurate than the "optimistic" CIs found in most papers that ignore worker-dependent variance.

3. The "Spam" Filter

How do you caught a worker who is just clicking "5" for everything? The system calculates a Correlation Coefficient () between a specific worker's scores and the global average. If a worker's correlation drops below 0.25, their data is nuked.

Experimental Results: The Blizzard Test

The authors put crowdMOS to the test against the Blizzard TTS Challenge. They compared 17 speech synthesizers. The results were startling:

  • Repeatability: Two separate runs of crowdMOS (different days, different workers) had a 0.99 correlation.
  • Accuracy: The scores matched paid UK undergraduates in a lab almost perfectly, except for two specific algorithms (T and W).

Blizzard Results Comparison

Insight: Workers with loudspeakers rated low-quality (narrowband) audio higher than lab users with headphones. This isn't a "failure" of the crowd—it actually reflects real-world usage. If your users are on laptops, the lab's "standardized" headphones might actually be giving you a false signal about what your users actually care about.

Critical Analysis & Conclusion

CrowdMOS is more than just a tool; it’s a shift in philosophy. It suggests that diversity of environment is a feature, not a bug.

Limitations:

  • Expert Tests: If you are testing high-fidelity codecs where you need 20 earbuds) will fail.
  • Screening Latency: You need a critical mass of scores before you can calculate the "global MOS" to filter out bad workers.

Future Work: This framework has already been extended to image quality and region-of-interest tasks, signaling a new era of "Human Computation" in signal processing.


References:

  1. ITU-T P.800 (The "Gold Standard")
  2. Amazon Mechanical Turk (The Platform)
  3. Open-source crowdMOS tools available via Microsoft Research.

Find Similar Papers

Try Our Examples

  • Find recent papers that utilize crowdsourcing for subjective quality assessment in multi-modal (video and audio) tasks and how they handle worker reliability.
  • Which original studies established the Two-Way Random Effects Model for MOS, and how has this paper simplified its implementation for missing data?
  • Search for research investigating the gap between "commodity hardware" (earphones/speakers) used in crowdsourcing and "high-end hardware" used in ITU-standardized laboratory settings.
Contents
CrowdMOS: Democratizing Subjective Audio Quality Assessment
1. TL;DR
2. The Motivation: Lab Guilt and Objective Failure
3. The Solution: CrowdMOS
3.1. 1. The Strategy for Human Intelligence Tasks (HITs)
3.2. 2. The Math: Modeling the Noise
3.3. 3. The "Spam" Filter
4. Experimental Results: The Blizzard Test
5. Critical Analysis & Conclusion