BGWO-Elite: Mastering Arabic Text Complexity via Swarm Intelligence and Elite Crossover
Feature selection using binary grey wolf optimizer with elite-based crossover for Arabic text classification
This paper proposes an enhanced Binary Grey Wolf Optimizer (BGWO) with an elite-based crossover mechanism to perform wrapper-based Feature Selection (FS) for Arabic text classification. By integrating the modified BGWO with an SVM classifier, the authors established a new State-of-the-Art (SOTA) performance on major Arabic corpora including Alwatan, Akhbar-Alkhaleej, and Al-jazeera-News.
Executive Summary
TL;DR: This paper introduces an advanced variant of the Grey Wolf Optimizer (GWO) designed to solve the high-dimensionality curse in Arabic text classification. By replacing the standard movement equations with an elite-based crossover, the authors created a wrapper-based Feature Selection (FS) method that outperforms BPSO and Bat Algorithms, achieving F-measures as high as 96% with minimal feature subsets.
Academic Positioning: This work acts as a technical refinement of metaheuristic search strategies. It bridges the gap between Swarm Intelligence (SI) and the unique linguistic challenges of the Arabic language (morphological richness and orthographic complexity), establishing a robust baseline for wrapper-based dimensionality reduction.
The "Curse of Dimensionality" in Arabic NLP
Arabic is characterized by a high degree of inflection—a single root can generate dozens of words. When documents are converted into a Vector Space Model (VSM) using TF-IDF, the resulting feature space is massive and sparse.
Current methods face a binary struggle:
- Filter Methods: Fast but "blind" to how features interact within a specific classifier.
- Standard Wrappers (BPSO, GWO): High accuracy potential, but they often fall into Local Optima (LO) because their search mechanisms become stagnant in high-dimensional binary spaces.
Methodology: The Evolution to BGWO2
The core innovation lies in the transition from BGWO1 (standard sigmoid-based binary GWO) to BGWO2 (Elite-based Crossover).
The Genetic-Swarm Hybrid
In a standard GWO, the "omega" wolves update their positions by averaging the coordinates of the Alpha, Beta, and Delta leaders. In a discrete binary space, "averaging" is often counter-intuitive.
The authors propose an Elite-based Crossover operator ():
- Instead of mathematical averaging, the new position of a wolf is determined by a random selection from the bits of the three leaders.
- This creates a "mutation-like" effect that allows the search agent to explore new regions of the feature space if the current leaders are stuck.
Figure 1: The proposed wrapper-based feature selection framework using BGWO.
Experimental Performance: SVM is the King of High Dimensions
The authors tested the optimizer across three classifiers: J48 (Decision Trees), Naive Bayes, and SVM.
Key Findings:
- SVM Superiority: Across Alwatan, Al-jazeera, and Akhbar-Alkhaleej datasets, the BGWO2 + SVM combination consistently yielded the highest F-measure.
- Feature Sparsity: As shown in the results, BGWO2 doesn't just improve accuracy; it selects the lowest number of features, effectively denoising the dataset.
Figure 2: Comparison of the number of selected features across different algorithms. Lower is better.
Quantifiable Success (Alwatan Dataset):
| Method | Precision | Recall | F-measure |
|---|---|---|---|
| BGWO2 (Ours) | 0.965 | 0.960 | 0.961 |
| BPSO | 0.956 | 0.951 | 0.952 |
| BBA | 0.960 | 0.956 | 0.957 |
Critical Insight: Why Does It Work?
The effectiveness of the BGWO2 + SVM pipeline stems from two factors:
- Search Stability: The elite-based crossover provides the "jump" needed to escape stagnation in the binary 0/1 space of feature inclusion/exclusion.
- Generalization: SVMs are inherently suited for text because they seek the maximal margin hyperplane, making them less sensitive to the noise that typically plagues Naive Bayes or Decision Trees in sparse spaces.
Conclusion and Future Outlook
The paper successfully demonstrates that meta-heuristic refinement is a viable path for optimizing Arabic NLP tasks. While the study focuses on traditional Machine Learning, the next logical step is applying BGWO2 to optimize the hyper-parameters and attention-head pruning of Large Language Models (LLMs) specialized for the Arabic dialect.
Takeaway for Practitioners: When your feature space exceeds 10,000 dimensions, don't just use a standard optimizer; look for variants that incorporate crossover or mutation mechanisms to maintain search diversity.
