Attack under Disguise: Weaponizing Trust in Crowdsourcing Systems
Aack under Disguise: An Intelligent Data Poisoning Aack Mechanism in Crowdsourcing
The paper proposes "Attack under Disguise," an intelligent data poisoning mechanism targeting the Dawid-Skene model in crowdsourcing. It utilizes a bi-level optimization framework to maximize label error while simultaneously inflating the calculated reliability (confusion matrix) of malicious workers to evade detection.
TL;DR
Researchers have uncovered a critical vulnerability in sophisticated crowdsourcing systems. While high-end models like the Dawid-Skene (DS) model are designed to filter out "bad" workers by measuring their reliability, this paper introduces an intelligent poisoning mechanism that does the opposite: it games the reliability estimation itself. By strategically agreeing with the majority on certain tasks, malicious workers "earn" high trust scores, which they then spend to overturn the results of target items.
The Evolution of the Poisoning Paradox
Crowdsourcing aggregates labels from non-experts to find the "ground truth."
- Naive Aggregation (Majority Voting): Vulnerable to simple "always lie" attacks.
- Sophisticated Aggregation (Dawid-Skene): Uses an EM algorithm to estimate a Confusion Matrix for each worker. If you lie too much, your weight drops to zero.
The Problem: An attacker with limited resources (only a few malicious workers) cannot win by brute force. If they use a naive attack, the DS model identifies them as "low quality" and ignores them. This creates a paradox: to be an effective attacker, you must first appear to be a high-quality worker.
Methodology: The Bi-Level Optimization Trap
The authors propose a bi-level optimization framework to find the "perfect lie."
1. The Objective Function
The attacker's utility function balances two goals:
- Attack Success: Changing the estimated label from the original to a flipped .
- Disguise: Maximizing the malicious workers' parameters and (reliability).
2. Overcoming the Discrete Hurdle
Since "flipping a label" is a discrete (0 or 1) event, the gradient is non-differentiable. The authors use a Sigmoid Approximation to smooth the objective function, allowing them to use gradient-based optimization.
3. Architecture of the Attack

The attack follows a two-step iterative cycle:
- Step 1 (Lower Level): Fix labels, run the Dawid-Skene EM to see how the system "perceives" the workers.
- Step 2 (Upper Level): Adjust labels using Projected Gradient Ascent to push the system toward the attacker's goal.
Experimental Proof: Stealth and Impact
The researchers tested the attack on real datasets like the Duchenne Smile dataset.
Breaking the System
The results are alarming. With just a small percentage of malicious workers, the "Intelligent Attack" (red line in the graph below) far outperforms "Inversion" (naive lying) and "Random" attacks.

The "Disguise" in Action
The most striking evidence is the distribution of reliability parameters. In the charts below, the malicious workers (red dots) are distributed in the top-right corner—meaning the DS model saw them as more reliable than the average human worker, even though they were sabotaging the data.

Critical Insight & Conclusion
This paper serves as a wake-up call for the security of human-in-the-loop systems.
- The Takeaway: Traditional "reliability estimation" assumes that errors are either random (spammers) or consistent (adversaries). It does not account for strategic behavior where an adversary purposefully builds "trust equity" to maximize the impact of a future lie.
- Limitations: The attack currently assumes the attacker can see normal workers' labels (full knowledge). While the paper shows the attack works with limited knowledge, real-world "blind" attacks would be harder to execute.
- Future Direction: Future crowdsourcing defense must look beyond simple accuracy and perhaps incorporate temporal consistency or behavioral patterns that are harder to spoof via gradient ascent.
