Creating an Argumentation Corpus: Bridging the Gap Between Legal Theory and Reality
Creating an argumentation corpus: do theories apply to real arguments?: a case study on the legal argumentation of the ECHR
This paper presents the development of a specialized legal argumentation corpus using 45 judgments from the European Court of Human Rights (ECHR). It integrates Pragma-dialectics and Argumentation Schemes to map real-world legal reasoning into structured datasets for machine learning.
TL;DR
This study tackles the difficult task of transforming complex legal judgments from the European Court of Human Rights (ECHR) into a structured argumentation corpus. By combining Pragma-dialectics and Walton’s Argumentation Schemes, the authors uncover deep-seated issues in how humans interpret legal logic and provide a roadmap for better AI-driven legal analysis.
Problem & Motivation: Why Theory Fails in Court
Most argumentation theories are born in the sterile environment of philosophy or logic. However, real legal documents are "anarchic" compared to mathematical proofs. The authors identify three primary hurdles:
- Ambiguity of Theory: Terms like "Gradualism" are poorly defined—is it a type of argument or just a structure?
- Implicit Information: Judges often leave critical premises unsaid (enthymemes), making it hard for annotators (or machines) to bridge the logical gap.
- Human Subjectivity: A lawyer’s personal "positivist" or "theorist" background significantly changes how they identify what constitutes an "argument" versus a "fact."
Methodology: The Core Framework
The researchers utilized a dual-layer framework:
1. Pragma-dialectical Structure
This layer defines how arguments are linked:
- Simple: A single defense.
- Multiple: Alternative defenses for one point.
- Compound: Chains of reasoning, either Coordinative (parallel) or Subordinative (serial).
2. Argumentation Schemes
Based on Douglas Walton’s work, they used 25 stereotypical patterns of reasoning, such as Argument from Precedent or Argument from Consequences.
Visual representation of simple, multiple, and compound argumentation structures used to classify the ECHR judgments.
Experiments & Results: Where Annotators Stumble
The experiment involved two sets of annotators: legal professionals and trained argumentation theorists. The data revealed a fascinating "Ontology of Law" problem. Annotators with strict positivist legal training often failed to recognize "Established Rules" as arguments, seeing them instead as static facts.
Key statistical insights:
- Most Common Scheme: Argument from Analogy (21.6%), followed by Argument from Established Rule (19.6%).
- Highest Error Rate: Argument from Established Rule had a 28% incorrect annotation rate, primarily because premises were scattered across multiple pages, making the "distance" too great for consistent human tracking.
| Argumentation Scheme | % Occurrences | % Incorrect Annotation |
|---|---|---|
| Argument from Analogy | 21.6% | 25.4% |
| Argument from Established Rule | 19.6% | 28.0% |
| Argument from Sign | 7.9% | 7.2% |
Depth Insight: The "Reported Argument" Breakthrough
One of the paper’s strongest contributions is the separation of Reported Arguments from Current Arguments. In the ECHR, the court first summarizes what the parties said (Reported) before giving its own verdict (Current). Treating these separately prevents the "noise" of rejected arguments from polluting the logic of the final decision—a critical insight for anyone building legal AI.
Critical Analysis & Conclusion
Takeaway
The paper proves that a "bottom-up" approach to annotation (identifying specific schemes immediately) is too difficult. Instead, a Taxonomic Hierarchy is required:
- Identify the general class (e.g., Causal, Abductive).
- Drill down into the specific scheme.
- Handle "Negative" versions of schemes explicitly.
Limitations
The sample size of 45 judgments is small by modern deep-learning standards, and the study does not account for how different languages (the ECHR operates in English and French) might shift the rhetorical nuances of the arguments.
Future Outlook
As we move toward LLM-based legal assistance, this paper reminds us that "Law is a discursive practice." AI shouldn't just summarize; it must map the dialectical tree of a case to be truly useful to judges and citizens alike.
