ALoDLM makes diffusion language models spend more compute on hard tokens by letting easy tokens commit early as context while unresolved tokens keep refining a hidden state through extra recurrent passes, reaching roughly 2.7× the throughput of a vLLM-served Qwen3 8B on GSM8K at matched accuracy.
If you want faster text generation than standard token-by-token decoding, Masked diffusion language model (dLM) are attractive: they fill in many masked positions in parallel instead of marching left to right. The catch, well-documented in prior work, is that same-size diffusion models consistently score lower than autoregressive ones on reasoning, math, and code benchmarks.
The authors diagnose a specific cause they call a computation-difficulty mismatch. In one denoising step, a standard DLM runs every masked position through the same fixed stack of Transformer layers. But some blanks are trivial (an article, a closing bracket) and others are hard (the next number in a calculation, a variable name that depends on earlier choices). Uniform compute wastes effort on easy slots and starves hard ones. Confidence-based decoding, used by models like LLaDA, helps a bit by deferring low-confidence tokens to a later step, but when a deferred token comes back it starts fresh from its mask embedding. The intermediate computation is thrown away.
That’s the gap ALoDLM targets: give hard tokens more compute within a denoising step, and let that extra compute accumulate rather than reset.
The core idea is a loop inside each denoising step. The model is split into three parts: a prelude (input embedding plus a few early layers), a recurrent core (a middle stack of Transformer layers that gets re-run), and a coda (the final layers plus the LM head). Each pass through the core refines the hidden state for every still-masked position. After each pass, two small heads read out the current state: the normal LM head produces a token distribution, and a new ExitGate head produces a halting probability.
The decoding rule per inner pass:
•
If a masked position’s predicted distribution has low entropy (below a threshold \u03c4), commit it: sample the token and freeze it as discrete context for later passes.
•
Otherwise, keep its evolving hidden state and run it through the core again.
•
Stop the inner loop when nothing is left, when the average cumulative halt probability crosses a threshold q, or at a maximum depth K (set to 4 in the paper).
Committed tokens matter because they change what the still-unresolved tokens are conditioning on. That is the mechanism that makes hard tokens benefit from waiting.
Training is the harder part. For a given masked example, each token’s exit depth is a discrete choice, and all those choices together form an exit schedule. The paper treats this schedule as a latent variable and derives a variational bound (Negative Evidence Lower Bound) that jointly trains the denoiser and the halting policy, with a geometric prior that gently prefers shallow exits. Because committing one token changes the context for the others, you can’t enumerate schedules analytically. The authors use a Score-function (REINFORCE) estimator (also called REINFORCE) with two variance-reduction tricks: a first-pass prediction loss as a baseline (a Control variate baseline), and a weighted auxiliary loss that supervises every intermediate readout a token passed through, not just the one where it exited.
h = prelude(x)
for s in range(1, K + 1):
h = recurrent_core(h) # refine latent state
r = coda(h)
probs, halt = lm_head(r), exit_gate(r)
commit = {i for i in unresolved if entropy(probs[i]) <= tau}
for i in commit:
h[i] = embed(sample(probs[i])) # discrete context
unresolved -= commit
if not unresolved or mean(cum_halt[unresolved]) >= q:
break
The authors convert Qwen3-1.7B and Qwen3-8B into ALoDLM by supervised fine-tuning on a 5B-token corpus, skipping the continued-pretraining stage that comparable converted DLMs like SDAR and WeDLM use.
•
On the 11-benchmark average spanning ARC, MMLU, MMLU-Pro, GPQA-Diamond, GSM8K, MATH-500, HumanEval, and MBPP, ALoDLM reaches 65.5 at 1.7B and 80.3 at 8B. Both beat the Qwen3 AR starting points (63.8 and 78.5) and every other diffusion baseline the authors tested.
•
Against the strongest diffusion baseline, WeDLM-8B, ALoDLM-8B is +5.2 points on the 11-benchmark average and wins 10 of 11 tasks. The code benchmarks show the largest gains.
•
On GSM8K with parallel decoding on a single NVIDIA B200, ALoDLM-8B reaches ~2.7× the throughput of Qwen3-8B served by vLLM at comparable or higher accuracy. At a matched 93.25% accuracy versus WeDLM-8B, ALoDLM delivers 8.5% more tokens per second and uses 13.6% fewer GFLOPs per generated token.
•
Analyses: raising the exit threshold q from 0.1 to 0.9 pushes the average loops-per-token from 1.6 to 2.34 and the 11-benchmark score from 77.9% to 79.1%, which the authors frame as test-time scaling by latent depth. Numerical tokens get the lowest first-pass halt probability, meaning the model learned, without any token-level labels, to keep refining digits longer than ordinary words.
•
Ablations: K=4 matches K=8 while training faster; K=2 plateaus lower. Putting the recurrent core in the middle layers beats putting it at the end. Removing the intermediate-supervision trick inflates denoiser gradient variance by up to 1.76× early in training.
The paper evaluates a whole system, so these gains blend architecture, objective, and the WeDLM-style conversion recipe. The ablations isolate the depth, placement, and variance-reduction choices, but not the exit-gate mechanism against a matched non-adaptive looped baseline.
•
If you are comparing diffusion LMs to AR models for latency-sensitive serving on long outputs, ALoDLM is a concrete existence proof that the usual quality gap is not inherent. The authors ship the 8B model and code, so this is testable rather than taken on faith.
•
The q and \u03c4 knobs give you a per-request quality-vs-throughput dial without retraining. On GSM8K, sweeping \u03c4 from 0.1 to 0.6 at fixed q=0.5 moved throughput from 279 to 508 tokens/s with roughly 1.5 points of accuracy cost. Worth measuring on your own workload before committing to a setting.
•
If you are building your own looped or early-exit model, the intermediate-supervision idea is cheap to try: supervise every pre-commit readout a token saw, weighted by the policy’s exit probabilities, rather than only the exit depth you sampled. It needs no extra forward passes.
•
Caveat on deployment: the throughput advantage over AR shrinks on short outputs and on domains unlike the fine-tuning data. The authors are explicit that time-to-first-token is worse than AR because prefill fills the KV cache at every recurrent depth.
•
All comparisons are on converted models starting from Qwen3; the paper does not train an ALoDLM from scratch, so results don’t speak to pretraining-scale behavior.
•
The 2.7× throughput claim is on GSM8K with specific (q, \u03c4) settings and a B200 GPU; the quality-efficiency frontier against WeDLM actually crosses, with WeDLM faster at the lowest-quality operating points.
•
The paper reports no ablation that swaps the adaptive gate for a fixed-depth loop of equal cost, so the share of the gain attributable to adaptivity specifically (versus deeper mid-layer recurrence plus the 5B-token SFT recipe) isn’t cleanly isolated.
•
Throughput is input-dependent by design: domains underrepresented in the SFT data may trigger more recurrent passes and fewer parallel commits, potentially erasing the speed advantage.