Beyond the Mother Tongue: Can Non-Native Listeners Accurately Judge Speech Quality?

Influence of Language Differences in Crowdsourcing Speech Quality Assessment Studies

2021-06-14
Rafael Zequeira Jiménez, Babak Naderi, Sebastian Möller
Summary
Problem
Method
Results
Takeaways
Abstract

This paper investigates the feasibility of using non-native crowd-workers for speech quality assessment, specifically evaluating German speech datasets with native English and Spanish listeners. Using the ITU-T Rec. P.808 framework, the authors demonstrate that while language mismatch introduces a slight overestimation bias, it maintains a high Pearson correlation (r > 0.86) with laboratory standards.

Executive Summary

TL;DR

This study challenges the traditional requirement that speech quality evaluators must be native speakers of the target language. By testing native English and Spanish speakers on German audio datasets, the researchers found that while non-native speakers tend to be more "forgiving" (overrating quality), their relative rankings are remarkably consistent with native experts.

Background Positioning

In the landscape of Quality of Experience (QoE) research, this work serves as a critical validation of ITU-T Rec. P.808 (crowdsourcing) against the legacy P.800 (laboratory) standards. It bridges the gap between strictly controlled phonetic studies and the pragmatic need for global, scalable telecommunication testing.

Problem & Motivation: The Native-Speaker Constraint

Traditionally, subjective speech quality testing was a "bottleneck" process. To get reliable Mean Opinion Scores (MOS), researchers needed to recruit native speakers into soundproof booths. With the rise of crowdsourcing platforms like Amazon Mechanical Turk, we can reach thousands of listeners instantly. However, a major doubt remained: If a listener doesn't understand the language, can they truly hear the distortion?

Previous research suggested that lack of semantic understanding could either make listeners more sensitive to noise (overestimating degradation) or make them miss subtle "unnatural" artifacts (overestimating quality). This paper aims to quantify that bias and see if it can be mathematically corrected.

Methodology: A Triple-Experiment Design

The authors utilized the SwissQual 501 dataset, which contains 50 distinct impairment conditions (ranging from network jitter to codec compression). They set up three distinct groups:

  1. E1 (Control): Native Germans (via Clickworker).
  2. E2/E3 (Experimental): Native English and Spanish speakers (via MTurk) with verified zero-knowledge of German.

The "Qualification" phase was particularly rigorous, involving trapping questions and content tests to ensure non-native speakers weren't secretly bilingual.

Model Architecture / Test Procedure Table 1: Comparison of Lab-MOS vs Crowdsourcing MOS for specific German conditions.

Experiments & Results: Mapping the Bias

The research revealed a fascinating "non-native lens":

  • The Overestimation Effect: Non-native listeners consistently gave higher scores to degraded speech. They often missed "robotic" sounds (conditions 26, 27) or subtle "electronic" artifacts that a native speaker would find jarring.
  • High Consistency: Despite the higher scores, the trend was the same. If a native speaker thought Condition A was worse than Condition B, the non-native speaker almost always agreed.

Experimental Results Comparison Figure 1: Scatterplot showing how non-native speakers (E2/E3) consistently "overrated" quality compared to native German listeners (E1).

To fix this, the authors used First-Order Mapping—a linear transformation that aligns the "generous" non-native scores with the native baseline. Post-mapping, the Root Mean Square Error (RMSE) dropped significantly, proving that the non-native data is not "noisy," just "offset."

Corrected Scores Figure 2: Results after applying first-order mapping, showing much tighter alignment between language groups.

Critical Analysis & Conclusion

Takeaway

The core contribution of this work is the verification that language comprehension is secondary to signal perception for most telecommunication impairments. This opens the door for companies to run "Global Listening Tests" without needing to source native speakers for every locale.

Limitations

  • Hardware Variance: The authors noted that laboratory listeners used professional gear, while crowd-workers used "daily life" headphones, which likely contributed to some missed artifacts.
  • Semantic Nuance: The study focused on signal degradations (codecs, noise). It remains to be seen if more subtle linguistic qualities (prosody, emotional resonance) can be judged by non-natives.

Future Outlook

This paves the way for a more universal objective model of speech quality. If we can reliably map Spanish ears to German speech, we are one step closer to a truly global standard for Quality of Experience that transcends linguistic borders.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize first-order or higher-order mapping to align subjective speech quality scores across diverse demographic groups in crowdsourcing.
  • What are the foundational papers defining ITU-T Rec. P.808, and how do they address the issue of listener reliability in unsupervised environments?
  • Explore research investigating whether the "overestimation bias" found in non-native speech assessment also applies to other audio domains like music quality or environmental noise evaluation.
Contents
Beyond the Mother Tongue: Can Non-Native Listeners Accurately Judge Speech Quality?
1. Executive Summary
1.1. TL;DR
1.2. Background Positioning
2. Problem & Motivation: The Native-Speaker Constraint
3. Methodology: A Triple-Experiment Design
4. Experiments & Results: Mapping the Bias
5. Critical Analysis & Conclusion
5.1. Takeaway
5.2. Limitations
5.3. Future Outlook