发表机构
Kuaishou Technology(快手科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SingDance是统一视频扩散框架,将语音表达设为语义角色,通过硬紧凑路由等技术实现组合式零样本歌舞视频生成,参数少且性能具竞争力。
AI 中文摘要
从参考图像、文本提示和音轨生成个性化舞蹈视频需要音乐条件下的身体动作,而歌舞任务还需满足第二个要求:可见主体必须同时清晰表达歌声。现有音乐条件方法主要关注编舞,而语音驱动模型通常假设可见主体产生输入语音,导致这种组合设置在很大程度上未被探索。我们提出SingDance,这是一个统一的视频扩散框架,将可控的语音表达定义为语义角色:可见主体要么是产生语音信号的源,要么是接收来自屏幕外表演者语音的听者。硬紧凑路由选择与任务相关的语音、音乐和角色条件,这些条件通过逐帧联合音频注入进行组合;源和听者共享相同的语音路径。训练采用不对称监督:屏幕内说话和精心策划的屏幕外对话响应视频建立角色控制,而器乐和仅歌曲舞蹈视频建立音乐条件下的身体动作。目标歌曲/源配置在训练期间从未出现。推理时,将源角色分配给歌曲,可组合分别学习的表达和歌曲条件舞蹈能力,实现组合式零样本歌舞。实验表明,该方法具有强的动作-节拍对齐和视觉保真度,能在保留音乐对齐身体动作的同时可靠切换语音表达配对,且与评估的最强语音驱动基线相比,生成时参数显著更少,唇同步性能极具竞争力。
英文摘要
Generating personalized dance videos from a reference image, text prompt, and audio track requires music-conditioned body motion. Singing-and-dancing adds a second requirement: the visible subject must also articulate the vocals. Existing music-conditioned methods focus primarily on choreography, while speech-driven models generally assume that the visible subject produces the input voice, leaving this combined setting largely underexplored. We introduce SingDance, a unified video diffusion framework that formulates controllable vocal articulation as a semantic role: the visible subject is either the source, who produces the vocal signal, or the listener, who receives it from an off-screen performer. Hard-compact routing selects task-relevant speech, music, and role conditions, which are composed through frame-wise joint audio injection; source and listener retain the same speech pathway. Training uses asymmetric supervision: on-screen speaking and curated off-screen conversational-response videos establish role control, while instrumental and song-based dancing-only videos establish music-conditioned body motion. The target Song/Source configuration is never observed during training. At inference, assigning the source role to a song composes separately learned articulation and song-conditioned dance capabilities, enabling compositional zero-shot singing-and-dancing. Experiments demonstrate strong motion--beat alignment and visual fidelity, reliable paired switching of vocal articulation while preserving music-aligned body motion, and highly competitive lip synchronization with substantially fewer generation-time parameters than the strongest speech-driven baseline evaluated.
Comments9 pages, 5 figures