VORACE: Demoting Hyper-parameter Tuning through Democratic Voting

Voting with Random Classifiers (VORACE): Theoretical and Experimental Analysis

2022-05-09
Cristina Cornelio, Michele Donini, A. Loreggia, M. Pini, Francesca Rossi
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces VORACE (VOting with RAndom ClassifiErs), an innovative ensemble technique that aggregates predictions from a pool of randomly generated, un-tuned classifiers using social choice voting rules. By treating classifiers as voters and class predictions as ranked preferences, VORACE achieves state-of-the-art performance across 23 UCI datasets, rivaling complex models like XGBoost and Random Forest.

TL;DR

Is it possible to achieve SOTA performance without spending days on hyper-parameter tuning? VORACE (VOting with RAndom ClassifiErs) says yes. By generating a "crowd" of random, un-tuned classifiers and aggregating their predictions through sophisticated voting rules (like Plurality or Borda), we can create an ensemble that is as accurate as XGBoost but far more sustainable and easier to deploy.

Problem & Motivation: The High Cost of the "Best" Model

In the current ML landscape, the "No Free Lunch" theorem forces researchers into a cycle of manual grid searches and domain-specific tweaking. Finding the optimal architecture or hyper-parameters for a specific dataset is computationally expensive and requires significant expertise.

The authors' insight is grounded in Social Choice Theory: instead of agonizing over finding the one perfect classifier, why not use the collective wisdom of many "mediocre" ones? If we treat each classifier as an independent agent (a voter) and the classes as candidates, we can use 200 years of political science and mathematical voting theory to pick the winner.

Methodology: The Core of VORACE

The VORACE workflow is elegantly simple yet theoretically grounded:

  1. Random Generation: Generate classifiers (Decision Trees, SVMs, or Neural Networks) with randomly sampled hyper-parameters (e.g., depth, hidden layers, kernels).
  2. Training: Train all classifiers on the same training set.
  3. Preference Extraction: For a new sample, each classifier outputs a probability vector. This vector is converted into a ranking (e.g., Class A > Class C > Class B).
  4. Voting Aggregation: A voting rule (Plurality, Borda, or even the NP-hard Kemeny rule) aggregates these rankings to find the winning class.

Model Architecture

The beauty of this approach is that it treats classifiers as Maximum Likelihood Estimators (MLE). Under certain noise models, voting rules are proven to be the optimal way to recover the "ground truth" ranking from noisy observations.

Theoretical Breakthrough: Beyond the Binary

While the Condorcet Jury Theorem tells us that a majority of independent voters is likely to be correct in binary choices, VORACE extends this to multi-class scenarios.

The authors provide a novel closed-form formula (Theorem 1) to calculate the probability of the ensemble being correct () based on the accuracy of individual classifiers (), the number of classes (), and the number of voters (). Crucially, they use Generating Functions to fix errors found in previous literature regarding Plurality voting math.

Experimental Results

The researchers tested VORACE on 23 UCI datasets, comparing it against heavy hitters like Random Forest and XGBoost.

MetricAverage ProfilePlurality (VORACE)XGBoost
Avg F1-Score0.86260.90060.8636* (binary only)

Performance Comparison

Key Findings:

  • Stabilization: Performance increases with the number of voters, plateauing around .
  • Superiority: VORACE consistently beats the best individual classifier in its own pool, proving that aggregation generates new value.
  • Efficiency: Plurality voting proved to be surprisingly robust, often performing as well as more complex rules like Copeland or Kemeny while being significantly faster.

Critical Analysis & Conclusion

Takeaway

VORACE is a "sustainable AI" win. It achieves high accuracy without the energy-intensive search for optimal hyper-parameters. It democratizes machine learning by allowing non-experts to build high-performance ensembles by simply "throwing random models at a voting booth."

Limitations & Future Work

The primary hurdle remains the independence assumption. In practice, classifiers trained on the same data often make correlated errors (making them "dependent voters"). While the authors address this via an "overlapping value" () analysis, future work could explore diversifying the data (e.g., bagging) to further decouple the voters. Additionally, applying VORACE to unstructured data like images or text remains an open frontier.

Ultimately, VORACE reminds us that in machine learning, as in democracy, the many are often smarter than the one.

Find Similar Papers

Try Our Examples

  • Search for recent papers that apply Social Choice Theory or committee selection voting rules to improve ensemble learning performance.
  • Identify the origin of the Condorcet Jury Theorem and how modern machine learning researchers have generalized it for non-binary classification tasks.
  • Which studies explore the "sustainable AI" aspect of avoiding hyper-parameter tuning through randomized model ensembles or similar low-resource techniques?
Contents
VORACE: Demoting Hyper-parameter Tuning through Democratic Voting
1. TL;DR
2. Problem & Motivation: The High Cost of the "Best" Model
3. Methodology: The Core of VORACE
4. Theoretical Breakthrough: Beyond the Binary
5. Experimental Results
6. Critical Analysis & Conclusion
6.1. Takeaway
6.2. Limitations & Future Work