GAN-AD: Tackling Healthcare Fraud with Unsupervised Generative Models
Unsupervised Anomaly Detection of Healthcare Providers Using Generative Adversarial Networks
This paper proposes GAN-AD, an unsupervised framework for healthcare provider anomaly detection using Generative Adversarial Networks. Applied to Medicare and private South African datasets, it leverages reconstruction errors to label potential fraud in the absence of ground truth, achieving up to 97.4% AUC in classification downstream.
TL;DR
With healthcare fraud siphoning off billions of dollars globally, the primary obstacle to AI intervention has been the lack of labeled fraud data. This paper introduces GAN-AD, an unsupervised framework that uses Generative Adversarial Networks to "learn" what normal provider behavior looks like. By identifying deviations from this norm, the model creates its own labels, which are then interpreted using SHAP to show exactly why a provider was flagged.
Context & Motivation: The "Label" Crisis
Detecting fraud in healthcare is not like detecting cats in photos; we rarely have a definitive "this is fraud" label. In South Africa and the US, billions are lost to "phantom billing" and "kickback schemes," but audit data is sparse and often private.
The authors' core insight is grounded in the "Normalcy" hypothesis: If a GAN is trained to generate "normal" practitioner claims data, any real claim that the GAN struggles to reconstruct—or that the Discriminator easily identifies as "not matching the known distribution"—is likely an anomaly.
Methodology: The Two-Step GAN-AD Framework
The researchers proposed a sophisticated workflow that bridges the gap between unsupervised deep learning and human-readable auditing:
- Unsupervised Label Generation: A GAN (Generator and Discriminator ) is trained on provider features (e.g., patient volume, HCPCS codes).
- Anomaly Scoring: They calculate an anomaly score based on the generator's reconstruction error and the discriminator's certainty.
- Supervised Refinement & Interpretation: Using the GAN-generated labels as "ground truth," traditional classifiers (XGBoost, Logistic Regression) are trained to identify feature importance via SHAP.

The Anomaly Score Physics
The score is defined as: This formula balances two things: how much the data looks like the training set () and how well the model can recreate that specific provider's behavior ().
Experimental Battleground: Medicare vs. Private Data
The authors tested GAN-AD on two distinct datasets:
- Medicare (Public): ~1 Million samples, 91 features.
- Private SA Dataset: Proprietary insurance data, 95 features.
Performance Benchmarks
The results confirm that while the framework is unsupervised, it can train highly accurate binary classifiers.
| Model | Medicare AUC | Private AUC | Private Sensitivity |
|---|---|---|---|
| Logistic Regression (LR) | 0.757 | 0.974 | 99.6% |
| XGBoost (XGB) | 0.747 | 0.903 | 99.4% |
The exceptional performance on the Private dataset (99.6% Sensitivity) suggests that in controlled, high-quality data environments, GAN-AD is nearly perfect at identifying suspicious providers.

Deep Insight: Beyond Detection to Interpretation
A "black box" flagging a doctor for fraud is a legal nightmare. To solve this, the authors used SHAP (SHapley Additive exPlanation).
As seen in the figure below, the model identifies specific behavioral triggers:
- Reporting Lag: A high delay between service and claim correlates strongly with fraud.
- Volume Anomalies: High numbers of unique beneficiaries for specific injury groups (HCPCS codes) push the prediction toward "anomalous."

Critical Analysis & Conclusion
Theoretical Contribution
This work shifts the paradigm from classification (needing labels) to distribution matching (needing only "normal" data). By using GANs as a label-generator, it maps a path for AI adoption in industries where data is "unlabeled" by nature.
Limitations
- Label Loop: There is a risk of a "circular logic" trap where the supervised model merely learns the biases of the GAN's reconstruction errors rather than true fraud.
- Imbalance: The Medicare results (Lower AUC) suggest that the GAN still struggles with highly sparse, high-cardinality public data compared to cleaner private datasets.
Final Thought
GAN-AD demonstrates that deep learning can be more than just a prediction engine; it can be an automated auditor. For the healthcare industry, this means moving from reactive investigation to proactive, AI-driven oversight.
