Get Started
Home
Topics
Search
Library
6 min read · Inference Optimization · Video Generation · Added Oct 9 · Paper published Oct 7, 2026

GRACE: Generation-aware latent compression for efficient video generation

Source: research paper via Hugging Face Daily Papers
Video diffusion inference is slow because every spatio-temporal token hits quadratic attention. GRACE keeps the pretrained VAE and DiT frozen, adds a residual latent, and supervises the encoder in the DiT’s own feature space — cutting tokens ~8× and latency 11.1× on Wan2.1-14B at matching VBench quality.
TL;DR
GRACE compresses a pretrained video Variational Autoencoder into a dual latent (frozen base plus learned residual) and aligns it inside the frozen Diffusion Transformer (DiT)'s own feature space, cutting token count by ~8× and inference latency by 11.1× on Wan2.1-14B while matching its VBench quality.
Why It Matters
Video diffusion models are slow because the transformer backbone runs attention over every spatio-temporal token the autoencoder produces, and attention cost grows quadratically with that count. Generating 81 frames at 480×832 with Wan2.1-14B takes roughly 14 minutes on an A100. The straightforward fix is to crank up the autoencoder’s compression ratio so there are fewer tokens to denoise.
That fix breaks two things. First, higher compression hurts reconstruction, which teams patch by widening the latent channels, but wider latents converge more slowly for the diffusion backbone. Second, swapping or heavily retraining the autoencoder shifts the latent distribution away from what the pretrained Diffusion Transformer (DiT) learned, so you either retrain the DiT from scratch (very expensive) or watch quality collapse. Prior work like DC-Gen retrofits a new high-compression autoencoder onto a pretrained DiT, but the authors show it drifts noticeably from the original pipeline’s appearance.
How It Works
The core bet: keep the pretrained autoencoder and DiT exactly where they already agree, and only add new machinery for the information that compression throws away.
Stage 1, dual latent. The frozen pretrained encoder runs on a spatially and temporally downsampled copy of the video, producing a base latent that lives in the exact distribution the DiT already understands. A new residual encoder sees the full-resolution video and emits extra channels carrying the detail and motion the downsampled input lost. The two are concatenated along the channel axis.
Stage 1, generation-aware alignment. Reconstruction loss alone pulls the latent away from the DiT’s preferred distribution. So the authors add a loss that noises both the compressed and original latents, pushes them through the frozen DiT, and matches their intermediate features via cosine similarity on the first 10 of 40 blocks. The autoencoder is tuned to produce latents the DiT finds easy to denoise, not just latents that decode sharply.
Stage 2, DiT adaptation with asymmetric denoising. The DiT gets widened input/output projections for the extra residual channels and is fine-tuned with LoRA. During sampling, the base latent is denoised at a lower noise level than the residual (offset δ=0.15), so the residual always builds on a cleaner anchor. Both are predicted in one forward pass, so no extra function evaluations.
# Stage 1: train residual encoder + decoder x_low = downsample(x, r_s=2, r_t=2) z_base = frozen_encoder(x_low) # pretrained space z_res = residual_encoder(x) # new channels z = concat([z_base, z_res], dim=C) loss = recon(decoder(z), x) + w * align(z, frozen_encoder(x), frozen_dit) # Stage 2 inference: base leads residual by delta at every step for u in schedule: tau_base = noise(max(u - 0.15, 0)) tau_res = noise(min(u, 1)) v_base, v_res = dit(z, [tau_base, tau_res], text)
What They Found
On VBench at 480×832×81 frames with Wan2.1-I2V-14B, GRACE lands within 0.02 of the uncompressed pretrained pipeline’s total score while running 11.1× faster (77.7s vs 863.2s per video on an A100). At 736×1280 the speedup grows to 15.5×. On text-to-video, GRACE actually beats the uncompressed pipeline on the semantic score (84.98 vs 78.70).
The ablation is the most informative part. The single-latent baseline (same token budget, no dual latent, no alignment) drops the T2V total to 82.12. Adding the dual latent recovers 1.43 points, generation-aware alignment adds 1.13 more, and the asymmetric denoising offset adds another 1.13.
Reconstruction and generation quality decouple visibly: Step-Video-VAE reconstructs 1.91 dB better than LTX-VAE yet scores 3.01 lower on VBench-I2V generation. GRACE-VAE itself reconstructs 1.13 dB worse than the single-latent baseline but generates 1.46 points higher. The authors read this as evidence that optimizing an autoencoder purely for reconstruction is the wrong objective when a pretrained DiT is downstream.
In a blind human study against DC-Gen, raters preferred GRACE on visual quality in 60–71% of comparisons across T2V and I2V.
What’s Useful
If you’re running a pretrained video diffusion stack and inference latency is the pain point, the GRACE recipe is worth studying because it needs no changes to the DiT architecture and no training from scratch. Total training cost reported is 38.5 H200 GPU-days per task (autoencoder + DiT adaptation), which is modest for a 14B-parameter pipeline.
The transferable lesson, even if you don’t adopt the full method: when you adapt or compress an autoencoder for a frozen downstream generator, supervise the autoencoder in the generator’s feature space, not just with pixel/perceptual losses. The ablation against V-JEPA 2.1 features shows that aligning to a general video foundation model helps less (semantic score 80.04) than aligning to the actual DiT you’ll generate with (82.55). The alignment target should be the model that will consume the latent.
The asymmetric denoising trick (denoise the “anchor” part of a composite latent ahead of the “detail” part) costs nothing at inference and is reusable in any pipeline that factors a latent into coarse and fine components.
Worth testing before committing: whether the gains hold on your base model. The paper validates one pipeline (Wan2.1-14B) and notes that applying GRACE to a new pretrained pipeline requires retraining both stages against that pipeline’s frozen components.
Caveats
Small objects suffer most under compression. Distant faces, text on signs, and thin structures get lost or broken even when overall scene and motion survive, visible in the paper’s own text-to-video samples. The effect weakens at 736×1280 because the same ratio leaves more tokens per frame.
GRACE-VAE trades reconstruction fidelity for generation quality, so it is not a drop-in replacement if you need the autoencoder for non-generative tasks. Adding more residual channels could help, but adapting a pretrained DiT to higher-dimensional latents is itself hard and the authors leave it to future work.
The method is tightly coupled to one pretrained pipeline at a time. The base latent comes from that pipeline’s encoder and the alignment loss is supervised by its DiT, so every new backbone costs another full training run. In image-to-video, human raters still preferred the uncompressed Wan2.1-14B over GRACE on visual quality (43.6% vs 23.7%), so the “matches the pretrained pipeline” claim is strongest on aggregate VBench totals rather than on every perceptual axis.
Topics
Inference Optimization
Video Generation
Computer Vision
Inference Optimization
Video Generation
Computer Vision
Up next in Inference Optimization
Long-WAM: Scaling the Context of World-Action Models
STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Inference Optimization140 episodes
Video Generation78 episodes
Computer Vision163 episodes