GEP-CV: Breaking the Unimodal Barrier of Expectation Propagation
Generalizing expectation propagation with mixtures of exponential family distributions and an application to Bayesian logistic regression
This paper introduces Generalized Expectation Propagation (GEP), a novel deterministic approximate inference framework that extends traditional EP by utilizing mixtures of exponential family distributions. To address convergence issues caused by noisy stochastic gradients, the authors further develop GEP-CV, which incorporates control variates for variance reduction, achieving superior performance in Bayesian logistic regression tasks.
TL;DR
Expectation Propagation (EP) has long been a staple of Bayesian inference, but its reliance on single-mode exponential family distributions often leaves it "blind" to the multi-modal reality of complex posteriors. This paper introduces Generalized Expectation Propagation (GEP), which adopts a mixture-model approach to capture multi-modality, and GEP-CV, which utilizes control variates to solve the stability issues inherent in stochastic optimization.
Academic Positioning
This work sits at the intersection of Deterministic Approximate Inference and Stochastic Optimization. It refines the "energy function" perspective of EP (similar to Power EP or BB-) but pushes the boundary of expressivity by moving beyond unimodal approximations.
Problem & Motivation: The "Averaging" Trap
Traditional EP attempts to match the moments of a target distribution. When the true posterior is multi-modal, a single Gaussian (the most common choice) will attempt to cover all modes, often resulting in a high-variance "average" that represents none of the modes accurately.
Furthermore, traditional EP has two major engineering flaws:
- Convergence: It is not guaranteed to converge, especially with non-smooth likelihoods.
- Memory: It requires storing "site parameters" for every single data point, a nightmare for Big Data.
The authors' insight is to treat the problem as a global optimization of the KL divergence where is a mixture distribution, and use stochastic methods to bypass the memory bottleneck.
Methodology: Mixture Models and Variance Control
1. From Single to Mixture
Instead of a single , GEP uses: This allows the approximation to "shape-shift" and fit multiple peaks in the probability landscape.
2. The GEP-CV Architecture
While stochastic gradients allow the model to scale, they are notoriously "noisy." To fix this, the authors introduce Control Variates (CV). By defining a function that is highly correlated with the gradient but has a known expectation, they can "subtract" the noise:
The algorithm iteratively draws samples from the mixture components and applies variance-reduced updates to both the mixing weights and the natural parameters .
Experimental Results: Stability and Performance
The authors tested GEP-CV on both synthetic three-modal data and several UCI real-world datasets for Bayesian Logistic Regression.
Quantitative Edge
GEP-CV consistently outperformed standard EP and Variational Inference (VI). In high-dimensional datasets like Madelon (501 dimensions), GEP-CV achieved significantly better log-likelihoods, proving that its mixture approach captures the complexity that unimodal methods miss.
Figure 1: RMSE comparison tracking iterations. GEP-CV (bottom line) shows the fastest and lowest error convergence.
The Power of Control Variates
The visual evidence of variance reduction is striking. When comparing the "Change of Approximating Mean" over time, GEP-CV exhibits significantly fewer oscillations than standard GEP.
Figure 2: Top row (GEP-CV) shows a much smoother convergence of parameters compared to the fluctuating bottom row (Standard GEP).
Critical Analysis & Conclusion
Takeaway
GEP-CV effectively solves the expressivity-stability trade-off. By using mixture models, it handles complex posteriors; by using control variates, it makes the training of those complex models feasible.
Limitations
- Computation: While memory-efficient, the sampling process for mixtures is computationally intensive compared to closed-form EP updates.
- Heuristics: The selection of the number of components () remains a hyperparameter. As seen in Table 4, more components aren't always better; was the "sweet spot" for 3-modal data, but larger increased optimization difficulty.
Future Outlook
The integration of Reparameterization Tricks (like those used in VAEs) could further enhance GEP, potentially allowing it to compete with modern Deep Bayesian frameworks in even higher-dimensional latent spaces.
