DMax: Smashing the Parallel Decoding Barrier for Diffusion LLMs

DMax: Aggressive Parallel Decoding for dLLMs

2026-04-01
Zigeng Chen, Gongfan Fang, Xinyin Ma, Ruonan Yu, Xinchao Wang
Summary
Problem
Method
Results
Takeaways
Abstract

DMax is a novel paradigm for Diffusion Language Models (dLLMs) that enables aggressive parallel decoding while maintaining generation quality. It outperforms the state-of-the-art LLaDA-2.0-mini by improving Tokens Per Forward (TPF) from 2.04 to 5.47 on GSM8K and achieving over 1,300 Tokens Per Second (TPS) on H200 GPUs.

TL;DR

DMax is a breakthrough in Diffusion Language Models (dLLMs) that solves the "one-way" decoding trap. By replacing binary mask-to-token transitions with a self-revising embedding-space transformation, it achieves a ~300% speedup in decoding parallelism without sacrificing accuracy. On H200 GPUs, it clocks in at a blistering 1,338 Tokens Per Second (TPS).

The Bottleneck: The "Point of No Return" in Masked Diffusion

Standard Diffusion LMs like LLaDA use a binary mask-to-token paradigm. During inference, the model picks high-confidence mask positions and replaces them with fixed tokens.

The problem? Error Accumulation. If the model makes a mistake in an early parallel step, those "wrong" tokens become fixed context for the rest of the sequence. There is no "undo" button. This forces researchers to use very conservative (slow) decoding to avoid semantic collapse.

Methodology: The DMax Recipe

DMax fixes this by turning the decoding process into a continuous "self-refinement" loop.

1. On-Policy Uniform Training (OPUT)

Standard Uniform Diffusion (UDLM) fails because it trains on random noise that looks nothing like actual language. OPUT instead uses On-Policy Rollouts:

  1. The model predicts tokens from a masked input.
  2. It then takes its own predictions (including the errors) and is trained to map them back to the clean ground truth.

This bridges the train-inference gap, teaching the model exactly how to fix its own typical mistakes.

DMax Training Overview

2. Soft Parallel Decoding (SPD)

Instead of committing to a hard token ID, DMax uses a Hybrid Embedding: By interpolating between the predicted token embedding and the mask embedding based on confidence (), the model preserves uncertainty. If the model is unsure, the "mask-like" nature of the embedding signals to subsequent layers that this position needs more refinement.

Soft Parallel Decoding Architecture

Performance: High Speed, No Compromise

The most impressive part of DMax is its Efficiency-Performance Trade-off. In the chart below (from the paper), you can see that while the original LLaDA-mini's accuracy falls off a cliff as speed (TPF) increases, DMax remains rock-solid.

  • GSM8K Accuracy: 92.1% (DMax) vs 92.6% (Base) — but DMax is 2.7x faster.
  • Throughput: Surpassed 1,300 TPS, making dLLMs truly competitive with highly optimized Autoregressive (AR) engines.

Experimental Results Comparison

Critical Insight: Why it Works

DMax's success hinges on the synergy between OPUT and SPD. You cannot have one without the other. SPD requires a model that "understands" mixed embeddings, and OPUT provides exactly that by training the model to be a "universal denoiser" for both masks and semi-correct tokens.

Conclusion & Future Outlook

DMax effectively removes the "speed-accuracy" trade-off that has plagued diffusion models in NLP. By enabling aggressive self-correction, it opens the door for dLLMs to be used in real-time reasoning and code generation where latency is critical.

Future Work: The authors suggest this approach could be scaled further via Mixture-of-Experts (MoE) or expanded to long-context (128k+) modeling where parallel decoding provides the highest marginal gains.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize on-policy sampling or reinforcement learning to improve the decoding trajectories of Diffusion Language Models.
  • Which original research first proposed the "Uniform Diffusion" training objective for text, and how does DMax's OPUT specifically solve its instability issues?
  • Explore if the Soft Parallel Decoding (SPD) concept of hybrid embeddings has been applied to vision-language-action (VLA) models for faster robot control generation.
Contents
DMax: Smashing the Parallel Decoding Barrier for Diffusion LLMs
1. TL;DR
2. The Bottleneck: The "Point of No Return" in Masked Diffusion
3. Methodology: The DMax Recipe
3.1. 1. On-Policy Uniform Training (OPUT)
3.2. 2. Soft Parallel Decoding (SPD)
4. Performance: High Speed, No Compromise
5. Critical Insight: Why it Works
6. Conclusion & Future Outlook