The Algorithmic Bench: Bridging Machine Learning Fairness and EU Law
Legal perspective on possible fairness measures – A legal discussion using the example of hiring decisions
This paper explores the alignment between mathematical fairness measures and EU anti-discrimination law, specifically within HR recruitment. It categorizes existing algorithmic fairness metrics (e.g., Independence, Separation, Counterfactual Fairness) and evaluates their legal feasibility under the German General Equal Treatment Act (GET) and EU directives.
TL;DR
As AI takes a seat at the HR table, the legal definition of "fairness" is undergoing a seismic shift. This paper argues that while many mathematical fairness metrics exist, most are legally incompatible with EU law. However, Counterfactual Fairness emerges as a potential gold standard, balancing the protection of individual freedom with the need for algorithmic accountability.
Contextualizing the Problem: Process vs. Result
Traditional anti-discrimination law (such as the German GET or EU Directives) is process-oriented. It asks: Was this specific individual treated differently based on a protected trait during the hiring process?
In contrast, AI is a "black box." We often cannot see the internal reasoning process. This forces a shift toward a result-oriented assessment: Is the outcome distribution fair? The authors highlight that this shift creates a legal vacuum, as current statutes are not designed to evaluate statistical parity over individual merit.
Methodology: The Taxonomy of Fairness
The authors categorize fairness measures into four distinct buckets, each with varying degrees of legal "fitness":
1. Group Fairness (Independence)
This is effectively a quota system. It ignores the "ground truth" (who is actually qualified) and mandates equal outcomes for groups.
- Legal Verdict: Highly problematic. It prioritizes social goals over individual personality development, potentially violating the core liberal values of EU primary law.
2. Group Fairness with Ground Truth (Separation & Sufficiency)
These measures (like Equalized Odds) depend on historical data.
- Visual Logic:

- The Trap: If the historical labels (the ground truth) are already biased (e.g., historical hiring patterns favoring men), these metrics simply "equalize the error," effectively cementing structural racism or sexism into the system.
3. Individual Fairness
"Similar people should be treated similarly."
- The Flaw: It assumes we can define a perfect "distance metric" for human similarity. In practice, this forces minorities to match the "majority prototype" to be considered "similar," which the authors argue is a decomposition of individuality.
4. Counterfactual Fairness (The Authors' Preference)
This asks: Would Bob have been hired if he were a woman, holding all other non-causal factors constant?
- Architecture: It uses a causal graph to "calculate out" the influence of protected features.

Critical Insight: Why Counterfactual Fairness Wins
The paper posits that Counterfactual Fairness is the only metric that aligns with the causal logic humans use to determine fairness. It doesn't rely on arbitrary quotas; instead, it enforces a standard where the protected attribute (like gender) has no causal influence on the final recommendation.
Experimental & Legal Reality
The authors demonstrate that these metrics are mutually exclusive. You cannot satisfy Independence, Separation, and Sufficiency simultaneously. This leads to a profound legal implication: Lawmakers, not just engineers, must choose which fairness they want.
Performance Metrics Comparison
The paper provides a rigorous reference for how these measures are calculated using the confusion matrix:

Deep Insight & Conclusion
The core takeaway is that transparency is the prerequisite for justice. While mathematical metrics are useful indicators for auditors, "Blackbox" models (like deep neural networks) pose a threat to individual rights because they cannot be easily cross-examined in court.
The Future Roadmap:
- Prefer Whitebox Models: Where possible, use interpretable models to preserve process-oriented legal review.
- Continuous Auditing: Fairness is not a "train and forget" metric; systems must be monitored for drift in real-world use.
- Specific Legislation: We need a new legal "discrimination" term that explicitly accounts for statistical evidence in the age of ADM.
Ultimately, the paper warns that if we let mathematical metrics alone define fairness, we risk reducing human potential to a preprogrammed percentage—a future that conflicts with the fundamental spirit of European law.
