CBEUS: Solving the Imbalance Bottleneck in Bankruptcy Prediction via Evolutionary Intelligence

Expert Systems With Applications

2025-01-01
Som Gupta
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces a Cluster-Based Evolutionary Undersampling (CBEUS) approach for bankruptcy prediction. It combines K-means clustering with Genetic Algorithms (GA) to optimize Artificial Neural Network (ANN) training sets by selectively removing majority-class instances.

TL;DR

Corporate bankruptcy prediction is a high-stakes task often crippled by "Data Imbalance"—where healthy firms vastly outnumber failing ones. This paper presents CBEUS (Cluster-Based Evolutionary Undersampling), a hybrid model that uses K-means clustering and Genetic Algorithms to surgically prune the majority class. By optimizing instance selection and neural network weights simultaneously, the authors achieved a massive leap in G-Mean (from 17.21% to 84.26%), ensuring that minority bankruptcy cases are no longer "ignored" by the AI.

The Imbalance Pain Point: Why Accuracy is a Lie

In financial datasets, bankruptcy is a rare event. If 95% of firms are healthy, a model can achieve 95% accuracy by simply predicting "No Bankruptcy" for everyone. This is the Accuracy Paradox.

  • Prior Work Failures: Random undersampling (RUS) loses valuable data, while oversampling (SMOTE) can introduce artificial noise.
  • The Insight: Not all non-bankrupt firms are equally useful for training. Some are "typical," while others are "noisy" or "outliers" that confuse the decision boundary. The authors propose that we should categorize the majority class first, then use evolution to find the best boundary for each sub-group.

Methodology: The GA-ANN Hybrid Architecture

The proposed CBEUS framework operates in three distinct phases:

  1. Structural Recognition (Clustering): The non-bankrupt firms are grouped using K-means. The "Silhouette statistic" is used to determine the optimal number of clusters (found to be ).
  2. Evolutionary Pruning: A Genetic Algorithm (GA) searches for specific distance thresholds () for each cluster. If a firm's distance from its cluster centroid exceeds the threshold, it is labeled as "noise" and removed.
  3. Simultaneous Optimization: Unlike traditional methods that treat sampling and modeling separately, the GA here optimizes both the selection rules and the ANN connection weights at the same time.

Model Architecture

The Chromosome Structure

The GA encodes a complex search space: Where represents the cluster thresholds and represents the neural network weights. By using the G-Mean as the fitness function, the model is forced to maximize the balance between Sensitivity (identifying failed firms) and Specificity (identifying healthy firms).

Experiments & SOTA Results

The researchers tested the model on 22,500 Korean manufacturing firms. The bankruptcy rate was a mere 5.9%, representing a significant "extreme imbalance" challenge.

Performance Comparison

The CBEUS method was compared against standard ANNs, Random Undersampling (RUS), and standard Evolutionary Undersampling (EUS).

MetricANN (None)ANN (RUS)GA-ANN (CBEUS)
Sensitivity2.96%69.63%87.41%
G-Mean17.21%78.99%84.26%
H-Measure42.5947.6056.16

Performance Visualization

The results show that CBEUS doesn't just improve accuracy—it fundamentally shifts the model's ability to "see" the minority class. The use of the H-measure further validates that the model is robust against varying misclassification costs.

Critical Analysis & Takeaways

The brilliance of this work lies in its Rule-Format Representation. Instead of a black-box sampler, the model generates an interpretable rule:

IF [Distance from Cluster_1 < 0.501] AND [Distance from Cluster_2 < 0.520] ... THEN Select Instance.

Limitations:

  • Computational Cost: GA is notoriously slow compared to gradient-based methods. Searching for both thresholds and weights simultaneously in a massive big data environment could lead to the "curse of dimensionality."
  • Clustering Sensitivity: The model's success is heavily reliant on the initial K-means quality.

Future Outlook: This research paves the way for "Self-Organizing" datasets where the AI actively participates in its own data cleaning. For financial institutions, this means more reliable early-warning systems and significantly lower risks of undetected corporate failures.

Find Similar Papers

Try Our Examples

  • Search for recent studies that utilize hybrid clustering and Metaheuristic algorithms (like Particle Swarm Optimization or Ant Colony Optimization) for financial distress prediction.
  • Which seminal paper first introduced the G-Mean as a standard metric for imbalanced classification, and how has its implementation evolved in modern deep learning frameworks?
  • Explore research applying evolutionary undersampling techniques to high-dimensional imbalanced datasets in fields such as medical diagnosis or cybersecurity fraud detection.
Contents
CBEUS: Solving the Imbalance Bottleneck in Bankruptcy Prediction via Evolutionary Intelligence
1. TL;DR
2. The Imbalance Pain Point: Why Accuracy is a Lie
3. Methodology: The GA-ANN Hybrid Architecture
3.1. The Chromosome Structure
4. Experiments & SOTA Results
4.1. Performance Comparison
5. Critical Analysis & Takeaways