arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15818cs.CV

FlowDance:基于音乐驱动的并行姿态与RGB流舞蹈视频生成

FlowDance: Music-Driven Dance Video Generation with Parallel Pose and RGB Streams

Genying Li, Boda Lin, Jiachen Li, Zijian Jia, Haojie Zheng, Yiming Wang, Shuchen Weng, Si Li

AI总结:

FlowDance是一种结合并行姿态与RGB流的音乐驱动舞蹈视频生成框架,通过时间步感知姿态注入和持久身份注入优化,构建了高分辨率野外舞蹈数据集,在相关任务上表现优异。

AI中文摘要:

音乐驱动的舞蹈视频合成旨在根据给定的音乐片段为参考人物制作动画,该任务极具挑战性,因为它要求模型同时学习音乐到动作的对应关系、保留身份的人体动画、时间一致性以及视觉逼真的视频生成。我们提出FlowDance,这是一种音乐驱动的舞蹈视频生成框架,通过并行姿态流和RGB流将显式运动建模与保留参考的视觉合成相结合。我们进一步引入了时间步感知姿态注入,以在去噪步骤中调整结构引导,以及持久身份注入,以在长视频中保留参考外观。为支持该任务,我们还构建了一个由流行度筛选的高分辨率野外舞蹈视频数据集,该数据集包含同步音乐、RGB视频、3D人体运动、相机参数以及投影的2D姿态标注。大量实验表明,FlowDance在舞蹈动作生成和音乐驱动的舞蹈视频合成中均取得了优异的性能。

英文摘要:

Music-driven dance video synthesis aims to animate a reference person according to a given music clip. The task is challenging because it requires a model to jointly learn music-to-motion correspondence, identity-preserving human animation, temporal coherence, and visually realistic video generation. We present FlowDance, a music-driven dance video generation framework that integrates explicit motion modeling with reference-preserving visual synthesis through parallel pose and RGB streams. We further introduce timestep-aware pose injection to adapt structural guidance across denoising steps and persistent identity injection to preserve the reference appearance over long video. To support this task, we further build a popularity-curated, high-resolution in-the-wild dance video dataset with synchronized music, RGB videos, 3D body motion, camera parameters, and projected 2D pose annotations. Extensive experiments show that FlowDance achieves strong performance in both dance motion generation and music-driven dance video synthesis.

↑