OISHO: Mastering MLP Training via Selfish Herds and Orthogonal Logic

Selfish herds optimization algorithm with orthogonal design and information update for training multi-layer perceptron neural network

2019-01-14
Ruxin Zhao, Yongli Wang, Peng Hu, Hamed Jelodar, Chi Yuan, Yanchao Li, Isma Masood, Mahdi Rabbani
Summary
Problem
Method
Results
Takeaways
Abstract

The paper introduces OISHO, an enhanced Selfish Herd Optimization algorithm incorporating orthogonal design and population information updates. It aims to optimize the connection weights and biases of Multi-Layer Perceptron (MLP) neural networks, achieving superior classification accuracy and convergence speed across 20 UCI datasets.

TL;DR

Training Multi-Layer Perceptrons (MLPs) is often a battle against local optima. This paper introduces OISHO, a hybrid meta-heuristic that blends Hamilton’s "Selfish Herd" theory with Orthogonal Experimental Design and dynamic information updates. By systematically exploring the parameter space and forcing population diversity, OISHO consistently outperforms standard optimizers like Whale Optimization (WOA) and Salp Swarm (SSA) in classification accuracy and stability.

Background: Beyond the Gradient Descent

While Back-Propagation (BP) is the industry standard, its reliance on gradients makes it fragile in "rugged" loss landscapes. Stochastic meta-heuristic algorithms offer a derivative-free alternative, but they often struggle with the "No Free Lunch" theorem—performing well on one dataset but failing on another. The authors identify that the original Selfish Herd Optimizer (SHO), which mimics prey-predator dynamics, specifically lacks the global search intensity required to tune hundreds of neural network weights simultaneously.

Methodology: The OISHO Architecture

The core of OISHO lies in how it fixes the "stagnation" problem of bio-inspired algorithms.

1. The Selfish Herd Logic

The algorithm divides the population into Prey and Predators. Prey seek the center of the herd for safety (global best), while marginal individuals are hunted and replaced. This creates a natural pressure for convergence.

2. Orthogonal Design (The Global Mutation)

Instead of random mutation, OISHO uses an Orthogonal Array (L9). This allows the algorithm to sample the search space using a statistically representative set of combinations. If a new "candidate" generated via this matrix is better than the global best, it's immediately adopted.

MLP Structure and Encoding Fig 1: The architecture of the MLP where weights and biases are encoded as the herd's position.

3. Dual Information Update

To prevent the population from huddling too closely (losing diversity), the authors implement a threshold-based update. Low-ranking individuals are "re-informed" using Gaussian distributions and random vectors from the population, ensuring the search doesn't get stuck in a single valley of the loss function.

Experimental Battleground

The authors tested OISHO against six major competitors (GG-GSA, GOA, GSO, SSA, WOA, and SOS) using 20 UCI datasets.

Key Results:

  • Accuracy Dominance: OISHO ranked 1st in 16 datasets and 2nd/3rd in the remaining 4.
  • Convergence Speed: In most tasks, such as the Blood and Liver Disorder datasets, OISHO exhibited a sharp "elbow" in its convergence curve, reaching lower MSE values significantly faster than rival algorithms.
  • Statistical Significance: Wilcoxon’s rank-sum tests resulted in p-values far below 0.05, proving OISHO's superiority isn't due to luck.

Performance Comparison Fig 2: Accuracy distribution across independent runs, showing OISHO's high stability compared to GG-GSA and others.

Critical Insight: Why Does It Work?

The magic isn't just in the bio-mimicry; it's in the mathematical rigor of the orthogonal design. By treating weight optimization as a "multi-factor experiment," the algorithm can find optimal directions that random movements would likely miss. The "Information Update" acts as a safety net, maintaining a entropy level in the population that standard SHO lacks.

Conclusion and Future Outlook

OISHO represents a powerful shift toward "hybrid" meta-heuristics—where biological intuition meets statistical DOE (Design of Experiments). While currently focused on MLPs, the authors suggest this framework could be extended to Recurrent Neural Networks (RNNs) and Radial Basis Function (RBF) networks. For researchers dealing with non-differentiable or highly complex objective functions, OISHO provides a robust, stable, and high-performance toolkit for parameter tuning.


Disclaimer: This analysis is based on the 2019 paper "Selfish herds optimization algorithm with orthogonal design and information update..." published in Springer Nature.

Find Similar Papers

Try Our Examples

  • Search for recent studies that integrate Orthogonal Experimental Design (OED) into other swarm intelligence algorithms like Harris Hawks Optimization or Grey Wolf Optimizer for neural network training.
  • Which paper first introduced the Selfish Herd Optimizer (SHO), and how does the predator-prey interaction mechanism in that original work differ from modern variants used in ANN optimization?
  • Explore the application of OISHO or similar meta-heuristic hybrid trainers in Deep Learning architectures beyond MLPs, such as Convolutional Neural Networks (CNNs) or Recurrent Neural Networks (RNNs).
Contents
OISHO: Mastering MLP Training via Selfish Herds and Orthogonal Logic
1. TL;DR
2. Background: Beyond the Gradient Descent
3. Methodology: The OISHO Architecture
3.1. 1. The Selfish Herd Logic
3.2. 2. Orthogonal Design (The Global Mutation)
3.3. 3. Dual Information Update
4. Experimental Battleground
4.1. Key Results:
5. Critical Insight: Why Does It Work?
6. Conclusion and Future Outlook