← Back to all articles
arXiv cs.AIOctober 7, 2026

Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

Excerpt

arXiv:2610.05608v1 Announce Type: cross Abstract: We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (I2AV) modes; a built-in super-resolution model raises the output resolution to Fu