Beyond Co-occurrence: Graph Embeddings for Precision Overtreatment Detection
Abuse detection in healthcare insurance with disease-treatment network embedding
2021-10-19
Summary
Problem
Method
Results
Takeaways
Abstract
This paper presents a three-stage network-based approach for detecting healthcare insurance abuse (overtreatment). It constructs a disease-treatment association network using Relative Risk (RR) and employs metapath2vec graph embedding for link prediction, achieving superior performance in identifying unnecessary medical treatments on real-world HIRA data.
## TL;DR
Overtreatment in healthcare is a multi-billion dollar drain on global insurance systems. This paper moves beyond traditional "scoring" models by building a **disease-treatment association network**. By using **metapath2vec** to embed these entities into a shared vector space, the authors can predict whether a prescription is medically "logical" for a given diagnosis, significantly outperforming traditional rule-based or co-occurrence-based detection systems.
## The Core Intuition: Why Simple "Co-occurrence" Fails
In the world of medical billing, a single claim often lists multiple diseases and multiple treatments. A naive machine learning model might assume that *every* treatment in that claim is linked to *every* disease mentioned.
This creates "noisy" data. For example, a patient might have a neck fracture (main disease) and a minor skin rash (sub-disease). If a doctor prescribes lumbar spine imaging, a naive model might incorrectly learn that neck fractures or rashes justify lower-back scans.
The authors' key insight is to use **Relative Risk (RR)** to filter these links. A connection only exists in their "Association Network" if the treatment is statistically significantly more likely to occur in the presence of that disease across the entire dataset.
## Methodology: The Three-Stage Framework
### 1. Constructing the Association Network
The network is heterogeneous, comprising Refined Diagnosis-Related Groups (RDRG), main diseases, sub-diseases, and three types of treatments: Procedures, Prescriptions, and Materials.

*As shown above, the association network (right) is much sparser and more clinically accurate than the dense, noisy co-occurrence network (left).*
### 2. Selecting the Best Graph Embedding
The paper evaluates several SOTA graph embedding methods (DeepWalk, node2vec, SDNE, etc.). However, because the network contains different *types* of nodes, **metapath2vec**—designed for heterogeneous information networks—emerged as the winner. It uses specific "meta-paths" (e.g., Treatment-Disease-Treatment) to capture the fact that two different drugs are similar if they are used to treat the same disease.
### 3. Overtreatment Detection via Link Prediction
Overtreatment is redefined as a **Link Prediction problem**. If the trained model predicts a "no-link" status between a treatment and *every* disease listed in a claim, that treatment is flagged as an abuse case.

## Experimental Results: Precision Matters
Testing on South Korea's HIRA (Health Insurance Review and Assessment) 2017 data, the model showed massive gains over the "Without Embedding" baseline:
* **Procedures**: Accuracy reached up to **97.6%**.
* **Prescriptions**: Accuracy peaked at **93.8%**.
The model proved particularly robust at identifying "unseen" patterns—cases where a specific disease-treatment pair was not in the training set, but the model could "infer" its legitimacy based on the similarity of the nodes in the latent embedding space.
## Critical Analysis & Conclusion
The value of this work lies in its **Inductive Bias**. By forcing the model to learn medical logic (the relationship between a diagnosis and an action) via a graph structure, it becomes much harder for fraudulent providers to "hide" abnormal billing patterns beneath complex claim filings.
**Limitations**: The model currently treats overtreatment as a binary "necessary vs. unnecessary" problem. It does not yet account for **over-dosage** (the treatment is correct, but the quantity is too high).
**Future Outlook**: Integrating this approach with external medical Knowledge Graphs (like DrugBank) could further enhance the model's clinical reasoning, allowing it to detect not just fraud, but also potentially harmful drug-drug interactions.
