Legal Documents Categorization by Compression: Beyond Feature Engineering
Legal documents categorization by compression
This paper explores automated text categorization of Italian legal excerpts, specifically addressing the "residual legislative power" problem. The authors introduce and compare character-based compression algorithms—AMDL, BCN, and NCD—against standard machine learning classifiers, achieving a peak accuracy of 75.71% using the Best-Compression Neighbor (BCN) approach.
TL;DR
Categorizing legal texts is a "hard" problem in AI due to linguistic complexity and the lack of large labeled datasets. This paper proposes a "feature-free" approach using data compression algorithms (like Gzip) to measure document similarity based on Kolmogorov complexity. In a study of Italian normative texts, the Best-Compression Neighbor (BCN) method significantly outperformed standard Support Vector Machines (SMO) and Decision Trees, proving that information theory can succeed where traditional NLP struggles.
Problem & Motivation: The "Residual Power" Grey Zone
In the Italian legal system, "residual legislative power" refers to all subjects not explicitly reserved for the State or concurrent jurisdiction. This "negative definition" creates a classification nightmare: identifying which subjects belong to the Regions requires analyzing hundreds of subjective, complex judgments from the Constitutional Court.
Traditional Text Categorization (TC) faces two major walls here:
- Semantic Overlap: Laws for "Industry" and "Business Incentives" are linguistically almost identical.
- The Small Data Problem: High-quality training sets for specific legal niches are rare. Traditional feature-based models (TF-IDF + SVM) require large amounts of data to overcome the noise of convoluted legal syntax.
Methodology: The Logic of "Zipping"
Instead of converting text into vectors of words (Bag-of-Words), the authors treat text as a string of information. They utilize the concept of Kolmogorov Complexity ()—the length of the shortest program that can produce a string .
Core Intuition
If document and document are similar, they contain similar patterns. If you concatenate them and compress the result (), the compressor will find those shared patterns and use fewer bits to represent the second document.
The authors investigated three specific protocols:
- AMDL (Approximate Minimum Description Length): Compressing a test file against a massive archive of all training documents in a class.
- BCN (Best-Compression Neighbor): A k-NN variant that finds the single training document that provides the best compression "boost" to the test document.
- NCD (Normalized Compression Distance): A mathematically rigorous metric that normalizes the distance between 0 (identical) and 1 (completely different).
Figure 1: Complex structure of legal judgments including "Object provisions" and "Parameters" which the algorithm must process.
Experiments & Results
The researchers tested these methods on a dataset of Italian object provisions spanning 7 categories (Agriculture, Tourism, Education, etc.).
| Algorithm | Accuracy |
|---|---|
| BCN (Compression) | 75.71% |
| NCD (Compression) | 64.29% |
| Naive Bayes (Standard) | 61.14% |
| SMO / SVM (Standard) | 50.86% |
| J48 / C4.5 (Standard) | 36.00% |
Key Findings:
- BCN Scalability: The BCN method was the clear winner. Interestingly, when the model was allowed to pick a "top 3" list of candidates, it was correct 97.14% of the time.
- Parameter-Free Advantage: Unlike SVMs, which require careful tuning of kernels and feature selection, compression techniques require zero preprocessing (no stemming, no stop-word removal).
- Robustness to Data Size: In a follow-up experiment on Italian literature (authorship attribution), the compression models scaled exceptionally well as document size increased, reaching 80.41% accuracy.
Figure 2: Accuracy comparison highlighting the superiority of compression modes (BCN/NCD) over standard ML (J48/SMO).
Critical Analysis & Conclusion
This paper serves as a vital reminder that in the era of high-dimensional feature engineering, the fundamental principles of Information Theory remain potent.
The Takeaway: Compression-based classification is an ideal "first-pass" filter. It handles the morphological variants of languages like Italian naturally and captures phrasal structures that simple word-counts miss.
Limitations: The current approach is computationally expensive for massive datasets because -NN requires comparing each test file to every training file. However, for specialized legal domains where the "training set" consists of a few hundred landmark cases, this approach is both more accurate and easier to deploy than a deep learning model.
Future Outlook: The authors envision an "extended architecture" where compression handles the initial candidate selection (shortlisting 2-3 classes), followed by a semantically-rich ontological tool to make the final legal distinction.
