MADM: Eliminating Discretization Bias in Diffusion Models with Metropolis Adjustments

Metropolis-Adjusted Diffusion Models

2026-05-01
Kevin H. Lam, Tyler Farghly, Christopher Williams, Jun Yang, Yee Whye Teh, Arnaud Doucet
Summary
Problem
Method
Results
Takeaways
Abstract

This paper introduces Metropolis-Adjusted Diffusion Models (MADM), a framework that integrates Metropolis-Hastings (MH) and Barker's accept-reject steps into the Predictor-Corrector diffusion sampling process. By utilizing a score-based line integral identity to estimate intractable density ratios, the authors provide the first exact correction method for diffusion models and a highly efficient Simpson's rule approximation to eliminate discretization bias.

TL;DR

Metropolis-Adjusted Diffusion Models (MADM) bridge the gap between MCMC theory and diffusion sampling. By replacing biased Langevin "corrector" steps with a novel, score-based accept-reject mechanism, MADM eliminates persistent discretization artifacts. It introduces an exact Bernoulli factory-based sampler and a hyper-efficient Simpson’s rule approximation that improves FID across standard benchmarks like ImageNet and FFHQ.

Motivation: The Hidden Bias in "Correction"

Diffusion models typically transform noise into data by simulating a reverse-time SDE. In the popular Predictor-Corrector (PC) framework, a predictor (like an ODE solver) moves the sample between noise levels, and a corrector (usually the Unadjusted Langevin Algorithm, or ULA) refines the sample.

However, there is a fundamental flaw: ULA is itself a biased sampler. Because it discretizes continuous Langevin dynamics, it never truly converges to the target distribution . In classical MCMC, we fix this using the Metropolis-Adjusted Langevin Algorithm (MALA). But MALA requires the ratio of densities , which is inaccessible in diffusion models—we only have the score .

Methodology: Score-Based Accept-Reject

The core insight of MADM is that while we don't know the density, we can calculate the Log-Density Ratio as a line integral of the score:

1. The Exact Approach: Two-Coin Bernoulli Factory

The authors propose the first exact Barker adjustment for diffusion. It uses a "Bernoulli factory"—a method to flip a coin with probability given flips of a coin with probability . By using a Poisson-truncated power series, they can decide to accept or reject a proposal based on the score function without ever approximating the integral.

2. The Practical Approach: Simpson’s 1/3 Rule

For large-scale image generation, the authors introduce a deterministic approximation using Simpson's Rule. By evaluating the score at the midpoint of a Langevin step, they achieve an error bound of . This is essentially "free" performance, requiring only one extra score evaluation per step.

Model Overview and Comparison Figure 1: Comparison of ODE Predictor, ULA Correction, and the proposed MADM Correction. Notice how MADM moves samples closer to the true data support.

Experiments and Performance

The researchers tested MADM on synthetic 2D Manifolds and high-resolution image datasets (CIFAR-10, FFHQ, ImageNet-64).

Key Findings:

  • Outlier Removal: On synthetic datasets (Spirals, Pinwheels), MADM successfully moved "hanging" outliers back onto the data manifold where ULA failed.
  • FID Gains: MADM provided systematic improvements in Fréchet Inception Distance (FID). Notably, on ImageNet, adding MADM to a simple Euler solver made it perform better than the significantly more expensive Heun solver.

Table of Results Table 1: MADM consistently lowers FID scores across various datasets and ODE solvers.

Critical Analysis & Conclusion

MADM is a mathematically elegant solution to a long-standing "shrugged-off" problem in diffusion sampling.

Pros:

  • Compatibility: Works with any pre-trained score-based model (Llama-gen, Stable Diffusion, etc.).
  • Theoretical Rigor: Provides the first exact MCMC adjustment for score-only targets.
  • Efficiency: The Simpson approximation is scalable for production environments.

Limitations:

  • The exact Bernoulli factory method is computationally expensive (random number of score hits).
  • It still assumes the learned score is accurate; if the neural network provides a poor score estimate, the "exact" adjustment will correctly sample a "wrong" distribution.

Takeaway: As generative models move toward high-precision applications (like protein design or architecture), the "good enough" discretization of ULA will no longer suffice. MADM provides the toolkit to turn diffusion models into precise statistical samplers.

Find Similar Papers

Try Our Examples

  • Find recent papers that apply Bernoulli factories or perfect sampling techniques to improve the accuracy of score-based generative models or SDE solvers.
  • Which original research introduced the Predictor-Corrector framework for Diffusion SDEs, and how does MADM specifically alter its mathematical convergence guarantees?
  • Explore studies that evaluate the robustness of Metropolis-Hastings adjustments when the learned score function has high estimation error or Lipschitz violations.
Contents
MADM: Eliminating Discretization Bias in Diffusion Models with Metropolis Adjustments
1. TL;DR
2. Motivation: The Hidden Bias in "Correction"
3. Methodology: Score-Based Accept-Reject
3.1. 1. The Exact Approach: Two-Coin Bernoulli Factory
3.2. 2. The Practical Approach: Simpson’s 1/3 Rule
4. Experiments and Performance
4.1. Key Findings:
5. Critical Analysis & Conclusion