Kandinsky 6.0 Video generates 5-second clips with synchronized 44 kHz audio and lip-sync by running a video stream and a from-scratch audio stream in parallel and wiring them together with bidirectional cross-attention inside every transformer block, with RL post-training cutting speech word error rate by about 47% on the 29B model.
If you want to generate a short video with matching sound, today’s open options force you to pick one modality and bolt the other on later. You generate the video, then run a separate text-to-audio model, then try to align them. Lip-sync usually breaks, impact sounds land on the wrong frames, and ambient audio has no relationship to what’s on screen. Closed systems like Veo 3.1 and Sora 2 do this jointly and well, but you can’t inspect them, fine-tune them, or run them locally.
The Kandinsky Lab team releases two open models under MIT license: a 3B Lite and a 29B Pro. Both do text-to-audio-video and image-to-audio-video, output 5-second clips, and upscale to Full-HD through a separate super-resolution stage. The baseline they most directly build on is their own prior Kandinsky 5.0 Video, which handled video only.
The core idea is to keep video and audio as two separate transformer stacks that talk to each other at every layer, rather than concatenating tokens into one stream. The video stack is initialized from the pretrained Kandinsky 5.0 video model. The audio stack is trained from scratch on 40M audio clips. Then they get fused.
Fusion happens through what the paper calls a dual-stream CrossDiT architecture. Inside each transformer block, the video tokens do self-attention and attend to the text prompt, then do cross-modal attention where video queries read from audio keys and values. The audio tokens do the symmetric thing. This is the mechanism that lets a hammer-strike frame pull the corresponding thud out of the audio stream, and lets lip movements line up with phonemes.
Position is encoded with Rotary Position Embedding (RoPE): video tokens get 3D positions (time, height, width), audio tokens get 1D temporal positions, normalized so the model handles variable frame rates. Training uses Flow matching with an MSE loss.
The pretraining is staged. First, the audio stream trains alone on text-to-audio. Separately, the video stream continues training on text-to-video using captions that also describe audio. Then both streams are connected through freshly initialized cross-attention layers and trained jointly on 7M paired audio-video clips, with image-conditioning mixed in at 25% during the final phase.
After pretraining comes a three-step post-training pipeline:
# Per-domain SFT, then merge, then RL, then distill
for domain in 11_domains:
ckpt[domain] = sft_stage1_visual(pretrained, domain) # 7k steps
ckpt[domain] = sft_stage2_sync_and_lipsync(ckpt[domain]) # 5k steps, adds face loss
sft_model = uniform_average(ckpt.values()) # model soup
rl_model = omninft_rl(sft_model, rewards=[hpsv3, clap, whisper_asr, lipsync, ...])
final = distill_pi_flow(rl_model); final = sim_ladd_adversarial(final) # to 10 NFE
RL uses OmniNFT adapted from a different open video model, with eight reward signals split across video, audio, and sync branches. A LoRA adapter is trained on top of the frozen SFT model, and gradients going from the video branch into audio keys and values are selectively scaled down to prevent one modality from clobbering the other. Distillation uses π-Flow to compress the sampling trajectory to 10 function evaluations, then an adversarial stage called Sim-LADD restores the texture detail that aggressive step-compression smooths away.
Separately, a text-free super-resolution model (Latent Upscaler plus an SR diffusion transformer, distilled to just two network evaluations) brings SD output up to Full-HD.
RL meaningfully improves speech intelligibility. On a held-out speech benchmark, word error rate for Pro drops from 0.235 (SFT) to 0.124 (RL), a 47% relative reduction, significant at p<0.001. Lite drops from 0.179 to 0.127. The gain holds on the public Harvard sentences set too. Note: WER is measured with Whisper-large-v3, which is also one of the reward models, so there’s some circularity. The Harvard-sentences result partially mitigates this.
On VABench, Pro leads the compared open models (Kandinsky Lite, LTX 2.5) in speech quality, audio aesthetics, text-video alignment, lip-sync, and visual realism. LTX 2.5 wins on VLM-judged alignment, expressiveness, and WER (measured here with a different ASR, so not comparable to the RL ablation above).
Human side-by-side evaluations are mixed and honest about it. Against the predecessor Kandinsky 5.0 Pro, the new model wins across all visual criteria. Against Kling 2.6, Veo 3.1 Fast, MiniMax H3, and Seedance 2.0, results split by criterion: Kandinsky Pro tends to lead on speech quality and sometimes on image-reference preservation and camera control; proprietary systems tend to lead on general visual quality and overall audio fidelity. MiniMax H3 and Seedance 2.0 are clearly stronger on visual criteria.
Distillation is nearly free. The distilled Pro model is preferred 51/49 against the full non-distilled version in head-to-head human eval, with no criterion reaching significance, while cutting inference to 10 NFE for the base model and 2 NFE for super-resolution.
Consumer GPUs work. Block offloading streams transformer blocks from CPU during the forward pass, keeping only two blocks resident. This runs Pro on 16 GB cards (RTX 5080 generates a 5s Full-HD clip in ~1546s) and 24 GB cards (RTX 4090, ~1247s). H100 does the same in ~402s.
•
If you’re building a product that needs joint audio-video and you can’t use closed APIs, this is one of a small number of fully-open 29B-scale options. Everything (weights, code, diffusers integration) is on GitHub under MIT.
•
If your use case is speech-heavy (talking heads, dialogue, dubbing), the paper’s strongest evidence lives there: the RL stage specifically targets lip-sync and ASR intelligibility, and the SBS evals show speech-quality parity or wins against commercial systems. Worth testing on your own prompts before committing.
•
If you need pure visual fidelity for cinematic content without heavy speech requirements, MiniMax H3 and Seedance 2.0 beat it in the paper’s own human evals. The authors are explicit about this.
•
The dual-stream + bidirectional cross-attention recipe (video stream pretrained, audio stream from scratch, cross-attention layers newly initialized and trained jointly) is a reusable pattern worth studying if you’re designing any two-modality generator where one modality already has a strong pretrained backbone.
•
The RL recipe (OmniNFT-style with modality-routed advantages, LoRA on the SFT base, per-block gradient-surgery schedule for cross-modal KV paths) is reportedly recalibrated for this backbone. The authors note that gradient-surgery zones from the prior LTX-2 work did not transfer directly, so expect to recalibrate if you adapt the method.
•
Clips are capped at 5 seconds and base generation is SD; Full-HD comes from the separate super-resolution stage, not from native high-resolution generation.
•
The RL WER improvement is partially circular: the main reward model (Whisper-large-v3) is also the evaluator. The Harvard-sentences replication helps but doesn’t fully resolve this.
•
Human-eval sample sizes are around 200 pairwise comparisons per criterion. Many reported differences do not reach statistical significance, and the paper flags this.
•
Russian-language training data (the Russian Cultural Code dataset of 1.2M images and 530K video scenes) is a deliberate design choice. Behavior on non-English, non-Russian prompts is not characterized.
•
Training-data sourcing is described in terms of filters and captioner models (Qwen variants, Gemma, VideoMAE for motion), but the ultimate provenance of the Kandinsky 5.0 video collection is not detailed in this paper. Treat licensing claims about generated outputs cautiously.