Can the Crowd Decode Privacy Policies? Scaling Legal Interpretation with ML
Crowdsourcing Annotations For Websites' Privacy Policies: Can It Really Work?
This paper evaluates the feasibility of crowdsourcing the annotation of website privacy policies. The researchers developed a specialized annotation tool and introduced a machine learning-based relevance model to highlight key paragraphs, successfully achieving over 95% accuracy when crowdworker agreement was high.
TL;DR
Privacy policies are the "unread terms" of the internet. This paper investigates whether we can use Amazon Mechanical Turk (the "Crowd") to annotate these dense legal documents. By combining a specialized annotation UI with a Machine Learning relevance model that highlights key paragraphs, the researchers achieved 95%+ accuracy and improved worker efficiency, proving that complex legal interpretation can indeed be crowdsourced.
Background: The "Notice and Choice" Failure
For decades, the "Notice and Choice" framework has failed because policies are written for legal compliance rather than human readability. Most users would need to spend an estimated 244 hours per year just to read the policies of the sites they visit. While automated NLP tools struggle with the inherent ambiguity of legal text, humans struggle with the sheer volume. This research explores the middle ground: Human-in-the-loop annotation supported by ML.
Methodology: Highlighting the Path to Accuracy
The researchers developed an interface that presents a policy alongside specific questions (e.g., "Does this site share health data?"). To solve the problem of workers missing info buried in long texts, they built a Relevance Model:
- Feature Engineering: They used normalized tf-idf of n-grams and custom Regular Expressions (Regex) curated by legal experts to identify data-practice-related phrases.
- Classification: A Logistic Regression model with L1 regularization (to prevent overfitting on a small dataset) predicted the probability of each paragraph's relevance to 9 specific privacy questions.
- The "Highlight" Intervention: In a between-subjects study, workers were shown either the full text (NOHIGH), or the top 5 (TOP05) or 10 (TOP10) predicted paragraphs highlighted in yellow.
Figure 1: The specialized tool showing paragraph highlighting and the navigation overview bar.
Experimental Results: Accuracy vs. Efficiency
The study compared 218 Turkers against 5 "Skilled Annotators" (Law/Policy graduate students).
1. The Power of Consensus
The researchers found that the crowd is remarkably accurate if you filter for agreement. By setting an 80% agreement threshold (at least 8 out of 10 Turkers must agree), the accuracy shot up to 96%. Crowdworkers rarely agreed on a "wrong" interpretation; if they couldn't agree, it was usually because the policy itself was genuinely ambiguous.
2. Efficiency Gains from ML
Does highlighting make workers lazy? Surprisingly, no.
- Time Savings: The TOP05 condition reduced the median task time by several minutes.
- Selection Integrity: Workers in the TOP05/TOP10 groups still manually selected text from non-highlighted areas when necessary, proving they didn't just blindly click on the yellow parts.
Figure 2: Annotation accuracy across different conditions. Highlighting did not negatively impact the quality of labels.
Deep Insights: Why Does This Work?
The success of this approach lies in Inductive Bias. By using expert-driven Regex to seed the ML model, the system guides the human eye to the most likely "crime scenes" in the document.
However, some practices remain "crowd-resistant." Data sharing with third parties (Q5-Q8) remains a major source of disagreement. This is likely because sharing clauses are often scattered: a policy might say "We don't share data" in section 1, but then list a dozen exceptions in section 10. This suggests that for highly distributed logic, we need even more sophisticated "context-aware" highlighting.
Summary & Future Outlook
This paper is a cornerstone for Usable Privacy. It proves that:
- Scalability is possible: We don't need a lawyer for every policy; 10 Turkers and an ML model can get 95% of the way there.
- UI Matters: Highlighting doesn't just save time; it improves the perceived self-efficacy of the workers, crucial for long-term labeling projects.
For future work, the industry is moving toward using these crowdsourced datasets to train Large Language Models (LLMs) to handle the "Unclear" cases, potentially bridging the gap where even consensus fails today.
Disclaimer: This analysis is based on Shomir Wilson et al. (2016). Original research was supported by the NSF and the Usable Privacy Policy Project.
