← Back to all articles
arXiv cs.AIAugust 18, 2026

Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer

Excerpt

arXiv:2608.15690v1 Announce Type: cross Abstract: Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output. We show that a base T2AV model can be turned into a voice-cloning model by adding a single zero-initialized linear layer on top of its audio backbone, fine-tuning for a comparatively short training schedule, and conditioning on a short reference recording at inference time. The ref