Expertise Screening: Turning the Crowd into Professional Image Quality Experts

Expertise screening in crowdsourcing image quality

2018-05-01
Vlad Hosu, Hanhe Lin, Dietmar Saupe
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces an advanced expertise screening framework for Crowdsourcing Image Quality Assessment (IQA). By implementing "gold standard" test questions focused on specific image degradations and refined User Rating-based Screening (URS), the authors achieved a 0.95 Spearman Rank Order Correlation (SROCC) between the screened crowd and freelance professional experts.

TL;DR

Is it possible to get professional-grade image quality ratings from unvetted crowd workers? This paper proves it is. By using a novel screening method based on artificial image degradations and sophisticated reliability checks, the authors achieved a 0.95 SROCC correlation between the general crowd and professional photographers, reducing data collection costs by over 80%.

Background: The Problem with Post-hoc Cleaning

Most crowdsourcing pipelines focus on "User Rating-based Screening" (URS)—essentially removing people who disagree with the majority. However, the authors argue that the "majority" isn't necessarily right if the majority is "naive." Expert photographers notice subtle artifacts (color fringing, lens blur, quantization) that the average user ignores. To build a gold-standard dataset like KonIQ-10k, we need the crowd to think like experts.

Methodology: The Expertise Filter

The core innovation lies in shifting from consistency screening to capability screening.

1. Artificial Degradation as "Gold Standard"

The researchers didn't just ask users to rate quality; they tested if users could identify why an image was bad. They created a dataset of 50 pristine images and manually injected 12 types of distortions:

  • Artifacts: JPEG compression, jitter, pixelation.
  • Blur: Lens blur, motion blur.
  • Contrast: Over-sharpening, exposure issues.
  • Colors: Over-saturation, color shifts.

Participants had to pass a "quiz mode" with 70-80% accuracy on these known distortions before their ratings were accepted.

2. Identifying "Uninformative" Users

Beyond random clicking, many workers are "honest but lazy," tending to stick to the middle of the scale. The authors introduced a metric to detect users with an unusually high frequency for a single answer. If , the user is discarded as uninformative.

Table 1: Experimental Setup and Costs

Experimental Insights: Does it Work?

The results were striking. The "Crowd 1" (C1) group, which used the degradation test questions, outperformed all other crowd groups in matching expert opinions.

The Repeatability Prediction

How do we know if the crowd is "as good as experts"? The authors used a clever extrapolation function to show that as the number of crowd judgments increases, their Mean Opinion Score (MOS) converges toward the repeatability limit of a group of 17 professionals.

Figure 2: Agreement Extrapolation The model accurately predicts how much the agreement improves as the group size grows.

The "Macro-Shot" Exception

An interesting finding emerged regarding Macro Photography. The crowd consistently rated macro shots (with shallow Depth of Field) lower than experts did.

  • The Crowd's View: "The background is blurry; therefore, the quality is low."
  • The Expert's View: "This is intentional lens blur for artistic focus; the technical quality is high."

Figure 5: Disagreement on Macro Shots Images where the crowd and experts diverged most due to intentional shallow depth of field.

When these 12 "optically complex" images were removed, the SROCC jumped to 0.945, nearly identical to the internal agreement of the experts themselves.

Critical Analysis & Conclusion

Takeaways

  • Capability Over Consistency: Screening users for their ability to find specific technical faults is far more effective than simply filtering out outliers.
  • Cost Efficiency: High-quality labels can be obtained at 1/6th the cost of hiring professionals if the screening pipeline is robust.
  • Domain Knowledge: Naive users still struggle with "intentional" degradation (like bokeh). Future screening should include "artistic intent" questions.

Limitations

The study relies on artificial degradations which are spatially uniform. Real-world images often have localized or depth-dependent distortions that might be harder for the crowd to identify even with this screening.

This work provides a blueprint for scalable, expert-level data collection for the next generation of AI-driven image enhancement and assessment tools.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize professional photographer benchmarks to evaluate the reliability of crowdsourced image quality or aesthetic datasets.
  • Which paper first introduced the concept of 'gold standard' questions in crowdsourcing, and how does the artificial degradation method in this paper iterate on that original concept?
  • Explore how the expertise screening techniques proposed here have been applied or adapted for subjective tasks in other domains like Video Quality Assessment (VQA) or Audio Quality Assessment.
Contents
Expertise Screening: Turning the Crowd into Professional Image Quality Experts
1. TL;DR
2. Background: The Problem with Post-hoc Cleaning
3. Methodology: The Expertise Filter
3.1. 1. Artificial Degradation as "Gold Standard"
3.2. 2. Identifying "Uninformative" Users
4. Experimental Insights: Does it Work?
4.1. The Repeatability Prediction
5. The "Macro-Shot" Exception
6. Critical Analysis & Conclusion
6.1. Takeaways
6.2. Limitations