Creating an Argumentation Corpus: Bridging the Gap Between Legal Theory and Reality

Creating an argumentation corpus: do theories apply to real arguments?: a case study on the legal argumentation of the ECHR

2009-06-08
Raquel Mochales, Aagje Ieven, A. Ieven
Summary
Problem
Method
Results
Takeaways
Abstract

This paper presents the development of a specialized legal argumentation corpus using 45 judgments from the European Court of Human Rights (ECHR). It integrates Pragma-dialectics and Argumentation Schemes to map real-world legal reasoning into structured datasets for machine learning.

TL;DR

This study tackles the difficult task of transforming complex legal judgments from the European Court of Human Rights (ECHR) into a structured argumentation corpus. By combining Pragma-dialectics and Walton’s Argumentation Schemes, the authors uncover deep-seated issues in how humans interpret legal logic and provide a roadmap for better AI-driven legal analysis.

Problem & Motivation: Why Theory Fails in Court

Most argumentation theories are born in the sterile environment of philosophy or logic. However, real legal documents are "anarchic" compared to mathematical proofs. The authors identify three primary hurdles:

  1. Ambiguity of Theory: Terms like "Gradualism" are poorly defined—is it a type of argument or just a structure?
  2. Implicit Information: Judges often leave critical premises unsaid (enthymemes), making it hard for annotators (or machines) to bridge the logical gap.
  3. Human Subjectivity: A lawyer’s personal "positivist" or "theorist" background significantly changes how they identify what constitutes an "argument" versus a "fact."

Methodology: The Core Framework

The researchers utilized a dual-layer framework:

1. Pragma-dialectical Structure

This layer defines how arguments are linked:

  • Simple: A single defense.
  • Multiple: Alternative defenses for one point.
  • Compound: Chains of reasoning, either Coordinative (parallel) or Subordinative (serial).

2. Argumentation Schemes

Based on Douglas Walton’s work, they used 25 stereotypical patterns of reasoning, such as Argument from Precedent or Argument from Consequences.

Model Architecture: Argumentation Structures Visual representation of simple, multiple, and compound argumentation structures used to classify the ECHR judgments.

Experiments & Results: Where Annotators Stumble

The experiment involved two sets of annotators: legal professionals and trained argumentation theorists. The data revealed a fascinating "Ontology of Law" problem. Annotators with strict positivist legal training often failed to recognize "Established Rules" as arguments, seeing them instead as static facts.

Key statistical insights:

  • Most Common Scheme: Argument from Analogy (21.6%), followed by Argument from Established Rule (19.6%).
  • Highest Error Rate: Argument from Established Rule had a 28% incorrect annotation rate, primarily because premises were scattered across multiple pages, making the "distance" too great for consistent human tracking.
Argumentation Scheme% Occurrences% Incorrect Annotation
Argument from Analogy21.6%25.4%
Argument from Established Rule19.6%28.0%
Argument from Sign7.9%7.2%

Depth Insight: The "Reported Argument" Breakthrough

One of the paper’s strongest contributions is the separation of Reported Arguments from Current Arguments. In the ECHR, the court first summarizes what the parties said (Reported) before giving its own verdict (Current). Treating these separately prevents the "noise" of rejected arguments from polluting the logic of the final decision—a critical insight for anyone building legal AI.

Critical Analysis & Conclusion

Takeaway

The paper proves that a "bottom-up" approach to annotation (identifying specific schemes immediately) is too difficult. Instead, a Taxonomic Hierarchy is required:

  1. Identify the general class (e.g., Causal, Abductive).
  2. Drill down into the specific scheme.
  3. Handle "Negative" versions of schemes explicitly.

Limitations

The sample size of 45 judgments is small by modern deep-learning standards, and the study does not account for how different languages (the ECHR operates in English and French) might shift the rhetorical nuances of the arguments.

Future Outlook

As we move toward LLM-based legal assistance, this paper reminds us that "Law is a discursive practice." AI shouldn't just summarize; it must map the dialectical tree of a case to be truly useful to judges and citizens alike.

Find Similar Papers

Try Our Examples

  • Search for recent legal argumentation corpora that utilize the ECHR dataset for Argument Mining tasks.
  • Which papers discuss the evolution of Walton’s Argumentation Schemes specifically applied to Automated Legal Reasoning?
  • Research how Large Language Models (LLMs) currently handle the distinction between "Reported Arguments" and "Current Arguments" in court judgments.
Contents
Creating an Argumentation Corpus: Bridging the Gap Between Legal Theory and Reality
1. TL;DR
2. Problem & Motivation: Why Theory Fails in Court
3. Methodology: The Core Framework
3.1. 1. Pragma-dialectical Structure
3.2. 2. Argumentation Schemes
4. Experiments & Results: Where Annotators Stumble
5. Depth Insight: The "Reported Argument" Breakthrough
6. Critical Analysis & Conclusion
6.1. Takeaway
6.2. Limitations
6.3. Future Outlook