Disaggregated Quantization (DQ) gives an LLM two different quantized versions of its own weights, one tuned for prompt processing and one for token generation, so each phase uses the number format its hardware bottleneck actually rewards.
When you serve an LLM, two very different things happen back to back. First, Prefill phase reads the whole prompt and builds the KV cache. This is a big matrix-matrix multiply, so the GPU is compute-bound: it wants low-precision arithmetic (like 4-bit multiplies) to go fast. Then Decode phase generates tokens one at a time. Each new token touches every weight once, so the GPU is memory-bound: it wants the weights to be as small as possible, even if the math itself stays in higher precision.
Today people pick one quantization scheme and use it for both phases. If you pick a hardware-native compute format like NVFP4, prefill flies but decode accuracy suffers more than it needs to (activation quantization hurts, and you’re not squeezing weights as small as you could). If you pick an aggressive weight-only format (say 2\u20133 bits per weight), decode is small and fast but prefill can’t use the fast 4-bit tensor cores. The authors’ observation: on Qwen 3 and Gemma 3, quantizing only decode to NVFP4 hurts reasoning benchmarks 2\u20134x more than quantizing only prefill, and the ordering flips on long-prompt short-answer tasks. So the two phases really do have different sensitivities, and one format for both is leaving accuracy on the table.
The intuition: treat prefill and decode as two collaborating sub-models that happen to share an attention architecture and a KV cache. Let each one pick its own weight representation and its own compute format. Train them jointly so they agree on what the KV cache should look like.
The training recipe is QADD (Quantization-Aware Distillation with Disaggregation). During supervised fine-tuning, every token already has a mask saying “this is prompt” vs “this is assistant response.” DQ reuses that mask to route each token through either the prefill pathway or the decode pathway of every linear layer in a single forward pass. The loss (KL against a frozen BF16 teacher) is only computed on response tokens, but gradients still reach the prefill weights because the response tokens attend to the prompt’s keys and values.
There are three flavors, in increasing ambition:
•
Format-disaggregated: same weights, but skip activation quantization during decode only. Free accuracy win, no extra storage.
•
Fully-disaggregated: train two separate weight checkpoints, one NVFP4 for prefill, one aggressively compressed (2\u20133 bit weight-only) for decode. Costs an extra checkpoint’s worth of storage.
•
Offloaded Disaggregated Prefill (ODP): for local single-GPU serving, keep only the decode weights resident and stream the prefill checkpoint from SSD block by block, overlapping the next block’s SSD read with the current block’s compute.
A fourth variant, prefillers, freezes an already-released aggressively-quantized decoder (e.g., a public 1-bit GGUF checkpoint) and trains only a matching NVFP4 prefill sidecar for it.
for token_pos in sequence:
if mask[token_pos] == PROMPT: # user turn
h = linear_prefill(h) # NVFP4 weights + NVFP4 activations
else: # assistant turn
h = linear_decode(h) # 2-3 bit weights, BF16 activations
# loss only on assistant positions; gradients flow to prefill via KV cache
loss = KL(teacher_logits, student_logits)[assistant_mask]
On Qwen 3 (0.6B\u20138B) and Gemma 3 (1B\u201312B), averaged across model sizes:
•
Format disaggregation (free intervention) lifts decode-heavy accuracy by roughly +1.9 to +4.5 points depending on format and family, with under 1.3-point movement on prefill-heavy tasks. No extra storage, decode is 2\u20133% faster because it skips activation quantization.
•
Full disaggregation at 2-bit decode is the headline: it beats non-disaggregated 2-bit by +7.4 to +10.7 points on decode-heavy tasks and by +5.3 to +10.5 points on prefill-heavy tasks. It even beats the slower weight-only 2-bit baseline by 4.5\u201312.5 points while keeping the fast NVFP4 prefill.
•
Prefillers on a public 1-bit Qwen3.8-27B GGUF: training an NVFP4 prefiller for the frozen IQ1_S decoder lifts MMLU-Pro by +32.5 points and MMMU-Pro by +35.3 points, more than doubling weight-only accuracy. The QADD training corpus contained no images, yet the visual-reasoning gain still transferred. Gains shrink as decode bitwidth rises; at 3 bits the prefiller can slightly hurt.
•
ODP latency: on Qwen3.8-27B in llama.cpp, time-to-first-token drops from 12.27s to 6.90s at 8K context (1.78x). Below ~4K context, SSD loading dominates and ODP is slower than the weight-only baseline. Above 16K, streaming overhead stays under 5%.
•
Large-scale PTQ check: applying just format disaggregation (no retraining) to eight models up to 2.8T parameters (including Kimi K3 and Nemotron 3) improved 11 of 13 model-benchmark combinations at 4-bit, with 6 statistically significant gains and zero significant regressions.
One wrinkle the authors surface: on MMLU-Pro, prefillers often shorten median response length but lengthen the mean (heavier tails). So “same decode checkpoint” does not mean “same tokens generated per query.”
•
If you serve an LLM with an already-quantized NVFP4 or W4A4 checkpoint, the cheapest intervention is format disaggregation: keep activations in BF16 during decode only. The paper shows a small accuracy bump and slightly faster decode with zero storage change, and it validated on models up to 2.8T with no retraining. Worth testing on your own quantized checkpoint.
•
If you have a very aggressively compressed decoder (1\u20132 bit) that you can’t retrain (e.g., a public GGUF), training just an NVFP4 prefiller is worth the effort. The paper’s released prefillers and code plus the llama.cpp fork with ODP and Hugging Face weights are the concrete artifacts to try first. At 3-bit and above, the prefiller can hurt, so measure before shipping.
•
If you’re on a single workstation-class device (their measurements are on DGX Spark) and your prompts are typically longer than ~4K tokens, ODP lets you get the accuracy of a two-checkpoint setup without paying for both in VRAM. Below 4K prompts, SSD loading dominates and it’s a loss.
•
If you serve mixture-of-experts models, the paper explicitly says ODP does not carry over well, because active-parameter ratios push the loading-vs-compute break-even out to impractical context lengths.
•
All accuracy claims are batch-one, single-turn. Multi-turn behavior is untested, and the authors flag a real concern: assistant tokens cached during decode carry decode-produced KV entries, but if a later turn rebuilds that cache through prefill, it will produce different representations for the same token history.
•
Latency numbers come from DGX Spark (GB10) with custom kernel integration into vLLM and llama.cpp. The prefill speedup ceilings (~1.5\u20131.7x over BF16) are diluted by attention, norms, and the unquantized LM head, and would look different on other hardware or with different attention backends.
•
The prefiller gains are largest at 1\u20132 bit decode, marginal at 2.5\u20133 bit, and slightly negative at 3-bit. This is a low-bit rescue tool, not a universal upgrade.
•
Full disaggregation stores two checkpoints. Without ODP, that’s roughly double the weight memory of the decode-only baseline; ODP recovers device memory but costs SSD bandwidth on every request’s first token.