Get Started
Home
Topics
Search
Library
7 min read · Audio/Speech · Multimodal · Added Oct 7 · Paper published Oct 4, 2026

Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

Source: research paper via Hugging Face Daily Papers
0:00 / 8:21
Open joint audio-video generation usually means bolting a TTS model onto a video model and praying lip-sync holds. Kandinsky 6.0 runs parallel video and audio transformers wired by bidirectional cross-attention at every block, with RL post-training cutting speech WER 47% on the 29B MIT-licensed model.
TL;DR
Kandinsky 6.0 Video generates 5-second clips with synchronized 44 kHz audio and lip-sync by running a video stream and a from-scratch audio stream in parallel and wiring them together with bidirectional cross-attention inside every transformer block, with RL post-training cutting speech word error rate by about 47% on the 29B model.
Why It Matters
If you want to generate a short video with matching sound, today’s open options force you to pick one modality and bolt the other on later. You generate the video, then run a separate text-to-audio model, then try to align them. Lip-sync usually breaks, impact sounds land on the wrong frames, and ambient audio has no relationship to what’s on screen. Closed systems like Veo 3.1 and Sora 2 do this jointly and well, but you can’t inspect them, fine-tune them, or run them locally.
The Kandinsky Lab team releases two open models under MIT license: a 3B Lite and a 29B Pro. Both do text-to-audio-video and image-to-audio-video, output 5-second clips, and upscale to Full-HD through a separate super-resolution stage. The baseline they most directly build on is their own prior Kandinsky 5.0 Video, which handled video only.
How It Works
The core idea is to keep video and audio as two separate transformer stacks that talk to each other at every layer, rather than concatenating tokens into one stream. The video stack is initialized from the pretrained Kandinsky 5.0 video model. The audio stack is trained from scratch on 40M audio clips. Then they get fused.
Fusion happens through what the paper calls a dual-stream CrossDiT architecture. Inside each transformer block, the video tokens do self-attention and attend to the text prompt, then do cross-modal attention where video queries read from audio keys and values. The audio tokens do the symmetric thing. This is the mechanism that lets a hammer-strike frame pull the corresponding thud out of the audio stream, and lets lip movements line up with phonemes.
Position is encoded with Rotary Position Embedding (RoPE): video tokens get 3D positions (time, height, width), audio tokens get 1D temporal positions, normalized so the model handles variable frame rates. Training uses Flow matching with an MSE loss.
The pretraining is staged. First, the audio stream trains alone on text-to-audio. Separately, the video stream continues training on text-to-video using captions that also describe audio. Then both streams are connected through freshly initialized cross-attention layers and trained jointly on 7M paired audio-video clips, with image-conditioning mixed in at 25% during the final phase.
After pretraining comes a three-step post-training pipeline:
# Per-domain SFT, then merge, then RL, then distill for domain in 11_domains: ckpt[domain] = sft_stage1_visual(pretrained, domain) # 7k steps ckpt[domain] = sft_stage2_sync_and_lipsync(ckpt[domain]) # 5k steps, adds face loss sft_model = uniform_average(ckpt.values()) # model soup rl_model = omninft_rl(sft_model, rewards=[hpsv3, clap, whisper_asr, lipsync, ...]) final = distill_pi_flow(rl_model); final = sim_ladd_adversarial(final) # to 10 NFE
RL uses OmniNFT adapted from a different open video model, with eight reward signals split across video, audio, and sync branches. A LoRA adapter is trained on top of the frozen SFT model, and gradients going from the video branch into audio keys and values are selectively scaled down to prevent one modality from clobbering the other. Distillation uses π-Flow to compress the sampling trajectory to 10 function evaluations, then an adversarial stage called Sim-LADD restores the texture detail that aggressive step-compression smooths away.
Separately, a text-free super-resolution model (Latent Upscaler plus an SR diffusion transformer, distilled to just two network evaluations) brings SD output up to Full-HD.
What They Found
RL meaningfully improves speech intelligibility. On a held-out speech benchmark, word error rate for Pro drops from 0.235 (SFT) to 0.124 (RL), a 47% relative reduction, significant at p<0.001. Lite drops from 0.179 to 0.127. The gain holds on the public Harvard sentences set too. Note: WER is measured with Whisper-large-v3, which is also one of the reward models, so there’s some circularity. The Harvard-sentences result partially mitigates this.
On VABench, Pro leads the compared open models (Kandinsky Lite, LTX 2.5) in speech quality, audio aesthetics, text-video alignment, lip-sync, and visual realism. LTX 2.5 wins on VLM-judged alignment, expressiveness, and WER (measured here with a different ASR, so not comparable to the RL ablation above).
Human side-by-side evaluations are mixed and honest about it. Against the predecessor Kandinsky 5.0 Pro, the new model wins across all visual criteria. Against Kling 2.6, Veo 3.1 Fast, MiniMax H3, and Seedance 2.0, results split by criterion: Kandinsky Pro tends to lead on speech quality and sometimes on image-reference preservation and camera control; proprietary systems tend to lead on general visual quality and overall audio fidelity. MiniMax H3 and Seedance 2.0 are clearly stronger on visual criteria.
Distillation is nearly free. The distilled Pro model is preferred 51/49 against the full non-distilled version in head-to-head human eval, with no criterion reaching significance, while cutting inference to 10 NFE for the base model and 2 NFE for super-resolution.
Consumer GPUs work. Block offloading streams transformer blocks from CPU during the forward pass, keeping only two blocks resident. This runs Pro on 16 GB cards (RTX 5080 generates a 5s Full-HD clip in ~1546s) and 24 GB cards (RTX 4090, ~1247s). H100 does the same in ~402s.
What’s Useful
•
If you’re building a product that needs joint audio-video and you can’t use closed APIs, this is one of a small number of fully-open 29B-scale options. Everything (weights, code, diffusers integration) is on GitHub under MIT.
•
If your use case is speech-heavy (talking heads, dialogue, dubbing), the paper’s strongest evidence lives there: the RL stage specifically targets lip-sync and ASR intelligibility, and the SBS evals show speech-quality parity or wins against commercial systems. Worth testing on your own prompts before committing.
•
If you need pure visual fidelity for cinematic content without heavy speech requirements, MiniMax H3 and Seedance 2.0 beat it in the paper’s own human evals. The authors are explicit about this.
•
The dual-stream + bidirectional cross-attention recipe (video stream pretrained, audio stream from scratch, cross-attention layers newly initialized and trained jointly) is a reusable pattern worth studying if you’re designing any two-modality generator where one modality already has a strong pretrained backbone.
•
The RL recipe (OmniNFT-style with modality-routed advantages, LoRA on the SFT base, per-block gradient-surgery schedule for cross-modal KV paths) is reportedly recalibrated for this backbone. The authors note that gradient-surgery zones from the prior LTX-2 work did not transfer directly, so expect to recalibrate if you adapt the method.
Caveats
•
Clips are capped at 5 seconds and base generation is SD; Full-HD comes from the separate super-resolution stage, not from native high-resolution generation.
•
The RL WER improvement is partially circular: the main reward model (Whisper-large-v3) is also the evaluator. The Harvard-sentences replication helps but doesn’t fully resolve this.
•
Human-eval sample sizes are around 200 pairwise comparisons per criterion. Many reported differences do not reach statistical significance, and the paper flags this.
•
Russian-language training data (the Russian Cultural Code dataset of 1.2M images and 530K video scenes) is a deliberate design choice. Behavior on non-English, non-Russian prompts is not characterized.
•
Training-data sourcing is described in terms of filters and captioner models (Qwen variants, Gemma, VideoMAE for motion), but the ultimate provenance of the Kandinsky 5.0 video collection is not detailed in this paper. Treat licensing claims about generated outputs cautiously.
Topics
Audio/Speech
Multimodal
Video Generation
Computer Vision
Audio/Speech
Multimodal
Video Generation
Computer Vision
Up next in Audio/Speech
VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models
Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations
Don't miss new content
Log in to follow topics and personalize your feed.
Related topics you might like
Multimodal127 episodes
Audio/Speech21 episodes
Video Generation72 episodes