Beyond Flat Data: Uncovering Hidden Discrimination via Semantic Ontologies
Classification Rule Mining Supported by Ontology for Discrimination Discovery
The paper introduces a novel framework for discrimination discovery that integrates ontology engineering with generalized classification rule mining. It leverages the legal methodology of "situation testing" within a semantic web context to identify unequal treatment in historical datasets, specifically demonstrated on the U.S. Harmonized Tariff Schedules (HTS) to uncover gender-based tax disparities.
TL;DR
Researchers from the University of Pisa have developed a framework that combines Ontology Engineering with Classification Rule Mining to detect systemic bias. By applying the legal principle of "situation testing" to a semantic hierarchy, they've demonstrated how to uncover complex discriminatory patterns in the U.S. Tariff system that traditional "flat" data mining methods would likely miss.
Problem & Motivation: The Limits of Flat Analysis
Most automated discrimination discovery tools treat datasets as simple tables. However, the real world is hierarchical. For example, in the U.S. Harmonized Tariff Schedules (HTS), a "tunic" is a type of "shirt," which is a type of "apparel."
Prior works often overlooked these relationships, leading to two major issues:
- Lack of Context: They couldn't distinguish between legally relevant attributes (PND) and protected characteristics (PD) within a nested hierarchy.
- Pattern Explosion: Without semantic grouping, analysts are buried under thousands of redundant rules.
The authors' insight was to use Ontologies as the backbone. This allows the system to compare "similar" individuals—not just those with identical values, but those who sit close to each other in a conceptual tree.
Methodology: Situation Testing at Scale
The core of the paper is the translation of the legal Situation Testing (or auditing) methodology into a computational logic.
1. Semantic Similarity
Instead of simple equality, the authors define similarity through Path Distance in the ontology. If two items belong to the same sub-class, their similarity is 1. If they share a parent, it decreases exponentially ().
2. The Discriminatory Indicator
For a given context (e.g., "Outerwear made of synthetic fiber"), the system looks for pairs of realizations where:
- The non-sensitive attributes are nearly identical (High Similarity).
- The protected attribute (e.g., Gender) is different.
- The outcome (e.g., Tariff rate) is significantly worse for the protected group.
3. Rule Extraction
The framework is implemented as a Protégé plugin, using SWRL (Semantic Web Rule Language) to query the knowledge base and extract generalized rules.
Fig 1: The proposed framework architecture, showing the flow from ontology population to rule generation.
Experiments: The HTS Case Study
The researchers applied this to the U.S. HTS dataset, specifically looking for gender bias in clothing taxes.
Key Findings:
- Footwear & Sleepwear: These categories showed the "steepest" distributions of discriminatory rules, indicating very specific and precise contexts where women's products are taxed higher than men's.
- Granularity Matters: The system found that "Shorts" had a higher confidence level for discrimination than the general "Pants" category. This "Multi-level" discovery allows analysts to pinpoint exactly where the bias originates.
Fig 2: Cumulative distribution of discriminatory rules. Note how different categories (Trousers vs. Pants) exhibit distinct bias patterns.
Critical Analysis & Conclusion
Takeaway
This work represents a significant step toward Semantic Data Mining. By using an ontology, the discovery process becomes more "human-centric" and aligned with legal reasoning. It successfully proves that gender-based tariffs in the U.S. aren't just random outliers but are embedded in specific material and form-based categories.
Limitations
- Scalability: While performance on the HTS dataset was fast, using complex SWRL queries on massive datasets (e.g., social media logs) might hit reasoning performance bottlenecks.
- Ontology Quality: The system is highly dependent on the "Domain Expert" to build a correct TBox. If the ontology is biased or incomplete, the discovery will be too.
Future Outlook
The authors suggest moving toward Dynamically Inferred Properties. Imagine a system that doesn't just look at static tags but uses reasoning to "understand" if a product is becoming more similar to another over time, adjusting its discrimination alerts in real-time.
