ÇORBA: Turning Legal Breaches into Security Requirements via the Crowd
C ¸ORBA: crowdsourcing to obtain requirements from regulations and breaches
The paper introduces ÇORBA, a novel crowdsourcing methodology designed to extract security and privacy requirements from regulatory texts (like HIPAA) and breach reports. By modeling requirements as "regulatory norms" (commitments, authorizations, prohibitions), the authors leverage human intelligence via Amazon Mechanical Turk to bridge the gap between abstract legal mandates and concrete technical safeguards.
TL;DR
Software developers often view regulations like HIPAA as "black boxes" of complex legalese. ÇORBA (Crowdsourcing to Obtain Requirements from Regulations and Breaches) provides a structured workflow to extract actionable security requirements from both abstract laws and messy real-world breach reports. By leveraging crowd workers on Amazon Mechanical Turk, the study demonstrates that with the right "question engineering," non-experts can outperform or match experts in identifying what went wrong and how to fix it.
The Motivation: Why Laws and Breaches Fail to Talk
The development of sociotechnical systems—where humans, organizations, and software interact—is governed by regulations. However, there are two major pain points the authors identify:
- Regulatory Abstruseness: Clauses are often vague (e.g., "Implement policies and procedures").
- Hidden Lessons: Thousands of breach reports are published every year, but the technical community rarely parses them for formal "Requirements Engineering" (RE) data.
Current NLP tools struggle with the nuance of legal "norms," and experts are too expensive. ÇORBA asks: Can we use the crowd to bridge this gap?
Methodology: The Architecture of a Norm
The "Secret Sauce" of ÇORBA is its formalization of requirements as Norms. A norm isn't just a rule; it’s a directed relationship consisting of:
- Subject: Who is responsible?
- Object: To whom are they responsible?
- Antecedent: Under what conditions does this apply?
- Consequent: What is the required action?
The Crowdsourcing Pipeline
The authors didn't just give workers a text box. They designed a two-phase process (Pilot and Final) that refined how we ask the crowd to "see" a law.

Key Innovation: Text Slimming
One of the most profound insights was that Breach Reports are noisy. They contain administrative fluff (e.g., "The Media was notified"). By "slimming" these reports down to core technical/causal sentences, the authors significantly boosted crowd accuracy.
Experiments & Results: "Verb-First" Instructions Matter
The study found that crowd workers initially struggled with Antecedents (the "When/If" conditions). In the final study, by forcing workers to start their answers with "If" or "When" and "Verb-first" for actions, the researchers saw a massive leap in response quality.

SOTA Findings:
- Effectiveness: Trained crowd workers achieved an average quality score of 0.81 (on a 0-1 scale).
- Experience vs. Performance: Surprisingly, workers with no prior experience in legal text performed just as well (0.82) as those with experience (0.80), proving that a well-designed methodology can empower the general population.
- Automation Potential: Using the crowd-curated labels, the authors trained a binary classifier (Doc2Vec + Logistic Model Trees) that could identify "useful" sentences in breach reports with 76.6% accuracy.
Critical Analysis: The Future of Regulatory RE
ÇORBA proves that security requirements shouldn't just come from the "top-down" (laws) but also from the "bottom-up" (breaches).
The Takeaway: If you want to build a compliant system, look at how others failed. The Limitation: While effective, manual evaluation of the crowd's responses is still a bottleneck. The authors suggest that the next step is using this high-quality crowd-labeled data to train LLMs or other automated tools to do the "slimming" and "norm extraction" at scale.
Conclusion
ÇORBA represents a shift from "compliance as a checkbox" to "compliance as a formal requirement." By structuring the collective intelligence of the crowd, the research turns a daunting pile of breach reports into a roadmap for a secure sociotechnical future.
Primary Contribution: A curated, publicly available dataset of 6,210 evaluated answers and 60 formal regulatory norms derived from HIPAA. Links: Dataset & Study Materials
