Preference Judgments in the Wild: Why "Ties" are Ruining Your Crowdsourced Data
Transitivity, Time Consumption, and Quality of Preference Judgments in Crowdsourcing
This paper investigates the reliability of preference-based relevance judgments in crowdsourcing, comparing Strict Preferences (A > B) and Weak Preferences (allowing ties, A ≈ B). It identifies that only strict preference judgments reliably maintain transitivity in a crowdsourced setting, achieving a quality level comparable to trained TREC assessors.
TL;DR
In the quest to evaluate search engine quality, asking crowdsourced workers to rank pairs of documents is better than asking them to assign grades. However, this study reveals a critical catch: forcing a choice (Strict Preference) is far superior to allowing a "tie" (Weak Preference). While strict preferences are 96% transitive and match expert quality, allowing ties disrupts logical consistency and increases the cognitive burden on workers.
The "Tie" Trap: Motivation and Insight
When evaluating if Document A is more relevant than Document B for a given query, it feels natural to allow workers to say "they are equally good." This is known as a Weak Preference. Intuitively, this should make the task easier.
The authors challenged this intuition. They noticed that while trained experts (like those at TREC) are highly consistent, crowdsourced workers are a "noisy" collective. Because different workers judge different pairs, the system's overall logic relies on Transitivity: if the crowd thinks and , it must conclude . If this property fails, the mathematical foundation for efficient ranking (reducing complexity to ) collapses.
Methodology: Measuring the Human Clock
The researchers didn't just look at the final labels; they "instrumented" the judging interface to measure:
- Reading Time: How long workers looked at the documents.
- Judgment Time: How long it took to click the final answer.
Figure 1: Comparison between isolated Graded Judgments and Pairwise Preference Judgments.
By aggregating at least three trusted judgments per pair, they tested whether the "wisdom of the crowd" remained logically sound across 12 different query topics from the ClueWeb12 dataset.
The Core Finding: Transitivity Breakdown
The results were striking. The logical consistency of the crowd depends almost entirely on the type of preference requested.
- Strict Preferences (A vs B): 96% Transitivity. It works. The crowd acts as a single, logical expert.
- Weak Preferences (A vs B vs Tie): 75% Transitivity. Even worse, when multiple "ties" were involved in a triple, the transitivity dropped to a dismal 32%.
Table 2: Transitivity holds strong for asymmetric (strict) relationships but fails in symmetric (tie) ones.
Why does this happen? The authors argue that a "tie" option introduces an ambiguous threshold. One worker’s "slight preference" is another worker’s "tie." This inconsistency creates "loops" in the data, making it impossible to produce a clean ranking.
Performance and Quality: Speed vs. Accuracy
One might think graded judgments (Standard IR practice) are faster. While a single graded judgment is faster because you only read one document, it provides much lower quality data.
- Agreement with Experts: Strict preferences achieved a Cohen’s κ of 0.530 compared to expert TREC judges, significantly higher than graded judgments (0.282).
- Efficiency: Strict preference judgments had the lowest "Judgment Time" (1.79s) compared to Weak (2.07s) and Graded (2.60s). Even though workers spend more time reading, they make decisions faster when the choice is binary.
Table 3: Strict preferences lead to the fastest decision-making once the documents are read.
Critical Analysis & Conclusion
This paper provides a vital lesson for AI and IR researchers: The UI design of your annotation task directly dictates the mathematical utility of your data.
Takeaway: If you want to replace expensive experts with crowdsourced workers, use Strict Preferences. Not only are they faster for the worker to process, but they also allow you to use sorting algorithms that assume transitivity, saving you thousands of dollars in redundant comparisons.
Limitations: The study focuses on Web Search (ClueWeb12). In highly subjective fields (like aesthetic judging or creative writing), the "tie" might be more prevalent, and forcing a choice could introduce different types of noise. However, for relevance, the "forced choice" is clearly the superior inductive bias for the system.
