GAN-Impute: Leveraging Generative Adversarial Networks for P2P Lending Risk Assessment
AI-Based Online P2P Lending Risk Assessment On Social Network Data With Missing Value
This paper introduces a Generative Adversarial Network (GAN) based approach for missing value imputation in online P2P lending risk assessment. By training a Generator to produce plausible data points and a Discriminator to distinguish them from complete records, the model effectively recreates social network data patterns to enhance credit prediction accuracy.
Executive Summary
TL;DR: This research tackles the pervasive issue of missing data in online P2P lending by introducing a GAN-based imputation framework. By training a generator to "fill in the blanks" and a discriminator to maintain data integrity, the method provides a robust pipeline for preparing messy social network data for high-precision risk assessment. It moves beyond simple statistical averages to capture the "hidden physics" of financial datasets.
Positioning: This work represents a transition from traditional statistical cleaning (Mean/Median) to AI-driven data synthesis in the FinTech domain, specifically targeting the high-sparsity nature of P2P lending features.
Problem & Motivation
In the world of online P2P lending, data is rarely pristine. Features such as "listing titles" or social network interactions often contain over 30% missing values.
The authors identify a critical flaw in current industry standards:
- Mean/Median Imputation: Reduces variability and underestimates variance, Leading to overconfident but inaccurate risk models.
- Deletion: Wastes valuable samples, which is unacceptable in specialized financial subsets.
- Linear Regression Imputation: Assumes a perfect linear correlation (1.0) between variables, which fails to capture the non-linear complexities of borrower behavior.
The motivation is clear: if we can simulate the distribution of the missing data using its relationship with existing features, we can build a much more resilient credit scoring model.
Methodology: The Adversarial Engine
The core innovation lies in applying the Generative Adversarial Network (GAN) architecture to tabular data imputation.
The Architecture
The workflow is divided into three distinct phases:
- Preparation: Complete data () is passed through a "Random Dropper" to simulate real-world sparsity ().
- Adversarial Training:
- Generator (G): Takes the incomplete record and a noise vector to produce a reconstructed record ().
- Discriminator (D): Evaluates whether the imputed values "fit" the patterns seen in complete records.
- Imputation & Classification: Once G is trained, it generates multiple candidates; the system uses a latent-space k-NN approach to select the most realistic values before feeding the data into a risk classifier.
Fig 1. Detailed GAN architecture for missing value imputation.
The Loss Function
Unlike standard GANs used for images, this model incorporates an MSE-weighted loss. This ensures that the generated data doesn't just "look" real but is also numerically close to the expected values in a financial context.
Experiments & Results
The authors tested their model using the Lending Club dataset from Kaggle, simulating a 20% data loss scenario.
SOTA Comparison & Stability
The experimental results show that the model is highly efficient:
- Convergence: Both training and testing MSE (Mean Square Error) show a sharp decline, stabilizing after approximately 1,000 epochs.
- Reliability: The generator successfully learns the internal correlations of the 18 columns (loan amount, interest rate, term, etc.), ensuring the "fake" data is indistinguishable from real records from the perspective of the discriminator.
Fig 2. Convergence of the MSE training loss over 5,000 epochs.
Critical Analysis & Conclusion
Takeaway
The study demonstrates that GANs are not just for generating "Deepfakes" or images; they are powerful tools for Data Augmentation and Cleaning. In financial risk assessment, where a single missing feature can lead to a wrong "Default" prediction, the ability of GANs to preserve the joint distribution of features is a game-changer.
Limitations
While effective, the paper relies on a relatively small sample size (1,000 records) and a simplistic "Random Missing" (MCAR) assumption. In real-world P2P lending, data is often "Missing Not At Random" (MNAR)—for example, borrowers with poor credit may intentionally hide certain information.
Future Outlook
The next step for this technology is its application in Unbalanced Data. In credit risk, defaults are rare events. Combining GAN imputation with SMOTE-like oversampling could yield a new generation of highly robust risk assessment engines.
