The Right Tool for The Job: Why Decision Trees Rule the UI Audit Space

The Right Tool for The Job? Assessing the Use of Artificial Intelligence for Identifying Administrative Errors

2021-06-09
Matthew M. Young, Johannes Himmelreich, Danylo Honcharov, Sucheta Soundarajan
Summary
Problem
Method
Results
Takeaways
Abstract

This paper investigates the application of Machine Learning (ML) to identify administrative errors in the US Unemployment Insurance (UI) system. The authors evaluate several models, concluding that CatBoost, a gradient-boosted decision tree framework, significantly outperforms deep learning architectures like TabNet and WideDeep in predicting payment errors.

TL;DR

Researchers from Syracuse University analyzed nearly two decades of US Unemployment Insurance (UI) data to see if AI can bridge the gap between efficiency and effectiveness. Their findings challenge the AI hype: CatBoost, a tree-based model, crushed modern Deep Learning architectures. More importantly, it offers the "explainability" required for democratic accountability, proving that in government, the most complex model isn't always the best.

The Public Values Conflict: A Hidden Seesaw

Every social safety net operates on a knife's edge. If you make it too easy to get benefits, you risk overpayment (inefficiency). If you make it too hard, you cause underpayment (ineffectiveness).

Historically, legislation has focused obsessively on the former—treating overpayment as "fraud" while ignoring the administrative wall that prevents eligible citizens from getting paid. The authors argue that AI could theoretically solve both simultaneously by identifying patterns of error before they become systemic burdens.

Methodology: Tree-Based Logic vs. Deep Learning Hype

The research utilized the DOL’s Benefits Accuracy Measurement (BAM) dataset, featuring 228 variables ranging from occupation codes to the date of the claim.

The team pitted traditional "shallow" learners against the titans of "Deep Learning":

  • CatBoost: A gradient boosting algorithm that excels at handling categorical data.
  • TabNet: Google's attempt to bring the power of attention mechanisms to tabular data.
  • WideDeep & DCN: Architectures designed to capture both linear and complex non-linear relationships.

Concept of Eligibility Overlap Figure 1: The trade-off between False Positives (Overpayment) and False Negatives (Underpayment) in decision-making.

The Results: A "Reality Check" for Deep Learning

The experimental results were stark. CatBoost didn't just win; it dominated.

ModelMacro F-scorePrecision (Underpayment)Precision (Overpayment)
CatBoost0.4590.6550.720
TabNet0.3130.3480.511
DCN0.3370.0000.565

The deep learning models (DCN and WideDeep) failed spectacularly at identifying underpayments, returning a precision of zero. This suggests that for tabular government data—which is often sparse, categorical, and heterogeneous—the inductive bias of decision trees is far more effective than the latent feature representations learned by neural networks.

CatBoost Feature Importance Figure 2: Key features driving CatBoost decisions, including occupation codes and claim timing.

Why CatBoost Won: Explainability as a Feature

In public administration, "The computer said no" is not a legal justification. Decisions must be scrutable.

  1. Interpretability: CatBoost allows administrators to rank feature importance. Knowing that "Occupation Code" or "Base Period Wage" are primary error drivers allows agencies to retrain staff or fix specific forms.
  2. Handling Noise: Public data is notoriously messy. Decision trees are naturally robust against the "noise" and outliers prevalent in UI claims.
  3. Efficiency: CatBoost achieved superior results with less computational overhead and no manual feature engineering.

Critical Insight & Future Outlook

The study highlights a critical "Goodness of Fit" problem. Many agencies are being sold high-cost Deep Learning "Black Box" solutions by vendors when simpler, tree-based models offer better performance and higher transparency.

Limitations: The authors used publicly accessible data. While robust, "live" internal government data (protected by privacy laws) might contain more granular signals that could potentially help Deep Learning models catch up—but for now, the Random Forest remains the king of the administrative jungle.

Takeaway for Practitioners: Don't let the "Neural" branding fool you. If you are working with tables and need to explain your results to a citizen or a judge, Gradient Boosting is the right tool for the job.

Find Similar Papers

Try Our Examples

  • Search for recent studies comparing Gradient Boosted Decision Trees (GBDT) vs. Deep Learning for tabular data in public administration or fraud detection tasks.
  • Which paper first proposed the Benefits Accuracy Measurement (BAM) auditing methodology, and how has the shift from manual to AI-driven auditing evolved since 2021?
  • Explore research that applies CatBoost or TabNet specifically to detect systemic "underpayment" or "denial of service" errors in other social safety net programs like SNAP or TANF.
Contents
The Right Tool for The Job: Why Decision Trees Rule the UI Audit Space
1. TL;DR
2. The Public Values Conflict: A Hidden Seesaw
3. Methodology: Tree-Based Logic vs. Deep Learning Hype
4. The Results: A "Reality Check" for Deep Learning
5. Why CatBoost Won: Explainability as a Feature
6. Critical Insight & Future Outlook