STEPQuant compresses the persistent recurrent state in gated linear-attention models to ~6 bits by giving long-lived state units more precision and fitting separate scales to key rows and value columns, matching FP32-state accuracy where naive INT6 collapses.
If you serve a hybrid model like Qwen3.8-27B or Kimi-Linear-48B-A3B-Instruct, you probably noticed something odd. These models replace most of their KV cache with a fixed-size recurrent state matrix per layer per head (via Gated DeltaNet (GDN) or Kimi Delta Attention), which sounds like a memory win. It is per request. But in a serving system like SGLang, every concurrent request needs its own persistent state in FP32, plus extra slots for caching and scheduling. At ~70 concurrent Qwen requests, the FP32 state pool exceeds the model’s BF16 weights. Short contexts don’t save you; concurrency is the bottleneck.
The obvious fix is to quantize the state the same way people quantize weights or KV caches. It fails badly. Naive uniform INT6 drops mean accuracy on Qwen from ~80% to ~45% on seven reasoning benchmarks, and INT4 falls near zero. The reason is specific to recurrent updates: each decoding step reads the quantized state, applies the Delta update, and writes a newly quantized state. Rounding errors feed back into the next step. Prior state-space quantization work like Q-Mamba applied dual-axis scaling to Mamba, but didn’t address error persistence in Delta-rule models.
The authors trace where quantization error does damage along two axes.
Temporal: some state units forget quickly, others don’t. Each head (Qwen) or key row (KDA) has a retention gate that controls how much of the old state carries forward. When that gate stays near 1, an error injected now persists for many steps. They verify this: heads with longer gate half-lives accumulate larger INT6 error (Spearman ρ ≈ 0.80 across 2,304 Qwen heads). So the fix is to spend more bits on long-lived units. They call this Lifetime-aware Bit Allocation: given an average bit budget, pick per-unit precisions from a candidate set to minimize distortion weighted by how long errors persist. A tiny fraction of especially risky units are kept in FP16 as pivots.
Spatial: the state matrix has outliers along BOTH key rows and value columns. Unlike activations in standard LLMs where outliers follow one axis, recurrent states show large-magnitude ridges on both. Worse, equal-magnitude errors in different key rows produce very different readout errors, because the readout multiplies the state by a query vector. The authors compute a row-impact score from calibration data (how much a row’s error affects the next readout) and then fit two scale vectors per head: a row scale that trades off row magnitude against row impact, and a column scale fit by minimizing impact-weighted reconstruction error.
for each decode step t:
S_prev = unpack(codes, row_scales, col_scales, fp16_pivots)
X = D_t @ S_prev + beta_t * k_t @ (v_t.T - k_t.T @ D_t @ S_prev)
y_t = X.T @ q_t # emit readout BEFORE requant
row_scales = sqrt(row_mag) / sqrt(row_impact_weight)
col_scales = fit_weighted_lsq(X, row_scales, impact_weights^2)
codes = quantize(X / (row_scales * col_scales), bits_per_unit)
Bit allocation and pivot selection run offline on 32 WikiText-2 segments; the row/column scale fit runs online per update, overlapped with later-layer compute on a separate CUDA stream.
Against the FP32-state baseline on seven long-generation reasoning tasks (including LiveCodeBench v6, AIME 2026, GPQA-Diamond, MATH-500) with BF16 weights:
•
Qwen3.8-27B at nominal 6 bits: STEPQuant averages 80.59% vs. FP32’s 80.60%; uniform INT6 gets 45.04%.
•
Qwen at 4 bits: STEPQuant 80.51% vs. uniform INT4 at 12.73%. The 4-bit variant essentially matches FP32 on Qwen.
•
Kimi at 6 bits: STEPQuant 61.47% vs. FP32 61.52%; at 4 bits it trails FP32 by about 3 points.
•
Short-generation tasks (MMLU, ARC-C, HellaSwag, etc.) show the same pattern, with 4-bit STEPQuant within 0.25 points of FP32 on both models.
The ablation separates contributions. On Qwen at 4 bits across three reasoning tasks: spatial fitting alone gets 73.95%, temporal allocation alone gets 12.87%, and combining both reaches 84.72% (FP32 is 84.61%). Protecting just 1.39% of Qwen heads as FP16 pivots adds 6.70 points at 4 bits. The adapted dual-axis component from Q-Mamba only reaches 7.64% at 4 bits on Qwen, showing that row-impact weighting matters, not just dual-axis scaling.
Two behavior findings worth noting. Uniform quantization causes overthinking: Kimi under INT4 generates ~63K tokens on AIME while scoring near zero, nearly hitting the 65K cap. STEPQuant keeps generation lengths close to FP32. And in SGLang at batch 512, 6-bit STEPQuant compresses recurrent-state memory 5.03×, cuts state-update time 2.91× on Qwen, and reduces total serving memory by 68.7% (Qwen W4) or 53.7% (Kimi W4).
•
If you serve a GDN or KDA hybrid model at high concurrency and your bottleneck is state-pool memory, this is directly applicable. The code is on GitHub, and the integration target is SGLang 0.5.12. The 6-bit setting is the safe recommendation; 4-bit is attractive on Qwen but costs real accuracy on Kimi for long generation.
•
If you’re evaluating state-space or linear-attention models, the lifetime-vs-error correlation (ρ ≈ 0.80) is a useful diagnostic on its own. Heads with retention gates near 1 are your quantization risk; measuring gate half-life on calibration data tells you where to spend precision before you commit to a scheme.
•
The row-impact score generalizes beyond this paper’s method. For any recurrent architecture where you read out via y = S^T q, equal-norm errors in different key rows have unequal readout impact. Worth testing as a sensitivity signal for pruning or mixed-precision work on similar architectures.
•
Combine with weight quantization. STEPQuant stacks cleanly on top of W4 AWQ weights with only ~0.05 point accuracy loss on Qwen at 6-bit state. As weights shrink, state becomes a larger fraction of serving memory, which is exactly when this helps most.
•
Evaluated on exactly two models (Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct) on four A800 GPUs with a fixed workload. Generalization to other GDN/KDA families or to dynamic serving traffic is not shown.
•
4-bit on Kimi loses meaningful accuracy on long reasoning (AIME drops from 68.33% to 59.17%). The authors flag this and attribute it partly to their lifetime weight being a gate-only approximation that ignores key-dependent transitions.
•
The comparison with concurrent work DAMP is not a controlled head-to-head: they used different generation settings and DAMP’s implementation wasn’t available. The retention-vs-bits comparison (STEPQuant 100.51% at 6.3 bits vs. DAMP 100.99% at 9.9 bits) relies on each method’s own FP32 baseline.
•
Calibration uses 32 segments of WikiText-2. The paper doesn’t study sensitivity to calibration domain, though a separate analysis shows gate-lifetime rankings transfer well across WikiText, C4, and LiveCodeBench text (Spearman > 0.98).