Automatic Legal Summarization: Beyond Mere Extraction to Rhetorical Structuring
Automatic Legal Text Summarisation: Experiments with Summary Structuring
This paper presents a machine learning-based extractive summarization system for legal texts, specifically UK House of Lords judgments. By employing Naïve Bayes and Maximum Entropy classifiers, the authors move beyond simple sentence selection to classify sentences based on their "rhetorical status," achieving performance significantly better than baseline methods in the legal domain.
TL;DR
Researchers from the University of Edinburgh have developed a system that summarizes complex UK House of Lords judgments by combining machine learning with rhetorical analysis. Unlike standard summarizers that just pick "important" sentences, this system identifies the role of each sentence—such as a statement of fact or a final ruling—to build structured, flexible summaries that are more useful for legal professionals.
The Challenge: Navigating the Labyrinth of "Legalese"
Legal judgments are notoriously difficult to summarize. A single case from the House of Lords can involve five different speeches, spans thousands of words, and uses convoluted language. Traditional extractive summarization—simply ranking sentences by frequency—often fails here because:
- Length and Complexity: Relevant information is buried under citations and procedural history.
- Multiple Perspectives: Different lords might agree on the outcome but for different legal reasons.
- Argumentative Structure: A summary needs to distinguish between the facts of the case and the final ruling (Disposal).
The authors' insight was that a summary shouldn't just be shorter; it should be structured according to its rhetorical intent.
Methodology: Teaching Machines the "Why" Behind the "What"
The core of the system is a two-step process: Classification and Ranking.
1. Rhetorical Classification
Sentences are assigned labels based on their function in the text. The labels include:
- FACT: Events leading to the case.
- PROCEEDINGS: Lower court history.
- BACKGROUND: Citations of law.
- FRAMING: The judge's own legal reasoning.
- DISPOSAL: The final ruling (the "bottom line").
2. Feature Engineering
To achieve this, the authors used a Maximum Entropy (ME) model fueled by several features:
- Location: Where does the sentence sit relative to the paragraph and the Lord’s speech?
- Cue Phrases: Automatically identified formulas like "It seems to me that..."
- Thematic Strength: Average tf*idf scores to find domain-significant terms.
Table 1: The rhetorical annotation scheme used to categorize legal discourse.
Experiments: Precision Matters
The study compared Naïve Bayes (NB) and Maximum Entropy (ME). While NB had higher recall in some areas, ME provided the high precision (71.4%) required for legal applications where accuracy is paramount.
Standard evaluation metrics (Precision/Recall) showed that the system performed exceptionally well on DISPOSAL sentences. This is critical because, for a lawyer, the most important part of a judgment is the actual decision.
Table 4: Cumulative performance of individual feature sets.
Designing a Better Summary
The paper introduces "Rhetorical Structuring." Instead of presenting sentences in the order they appear in the original text, the system can:
- Group sentences by the Lord who spoke them.
- Within each speech, present FACTS first, followed by FRAMING, and ending with the DISPOSAL.
This mirrors how a human lawyer would explain a case, providing a logical narrative flow that raw extraction lacks.
Critical Insight & Future Outlook
The heavy lifting in this paper isn't just the machine learning—it's the corpus annotation. By manually aligning legal abstracts with original documents, the researchers created a "gold standard" that reveals a hard truth: human summaries are highly abstract.
Limitations:
- The transition between extracted sentences can still feel jarring ("Discourse Smoothing").
- The system struggles to "blend" information across Lord's speeches as human writers do.
Future Work: The authors propose user studies to see how these automated summaries actually save time for legal researchers in the real world. As we move toward Large Language Models (LLMs), the rhetorical labels pioneered here remain vital for ensuring AI stays grounded in the specific logic of legal argumentation.
Takeaway
Summarization isn't just about compression; it's about understanding role and status. By identifying the "Disposal" and "Framing" of a legal text, we can turn a wall of text into a structured map of a court's decision.
