arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09581cs.CVcs.SD

万舞者:一种用于分钟级连贯音乐到舞蹈生成的分层框架

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

Mingyang Huang, Peng Zhang, Li Hu, Guangyuan Wang, Ruoshi Zhang, Yi Lu, Gang Cheng, Bang Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对从音乐生成舞蹈视频的挑战,提出万舞者分层框架,解耦过程为全局关键帧规划和局部时间细化,利用音乐上下文确保连贯,通过动态帧率自适应等创新,突破时长限制,在多舞蹈类型上表现出色,达新的技术水平。

中文摘要 AI 辅助

直接从音乐生成长时间、高清且节奏同步的舞蹈视频仍然是一项重大挑战,主要是由于当前扩散模型的时间限制,通常在超过20秒时就会失败。现有方法存在时间漂移、身份不一致和重复运动模式等问题。为此,我们提出了一种用于分钟级连贯音乐到舞蹈生成的新型分层框架。该方法将过程解耦为全局关键帧规划和局部时间细化,利用全轨道音乐上下文确保长程连贯性。关键创新包括通过时间映射的RoPE嵌入进行动态帧率自适应以实现精确对齐、基于光流的损失函数增强运动连续性以及运动速度控制以在快速运动中保留高保真细节。大量实验表明,我们的框架突破了传统时长限制,生成稳定的720p/30fps、超过一分钟的视频,具有卓越的时间稳定性。此外,该模型在五种不同舞蹈类型上表现出强大的通用性,以音频和文本提示为条件,在连贯的长格式舞蹈视频合成方面建立了新的技术水平。

英文摘要

Generating long-duration, high-definition, and rhythmically synchronized dance videos directly from music remains a significant challenge, primarily due to the temporal constraints of current diffusion models, which typically fail beyond 20 seconds. Existing approaches, whether they rely on intermediate 3D skeletons or on end-to-end video synthesis, suffer from temporal drift, identity inconsistency, and repetitive motion patterns when extended to longer horizons. To address these limitations, we propose a novel hierarchical framework for minute-scale coherent music-to-dance generation. Our method decouples the process into global keyframe planning and local temporal refinement, leveraging full-track musical context to ensure long-range coherence. Key innovations include dynamic frame rate adaptation via time-mapped RoPE embeddings for precise alignment, an optical-flow-based loss function to enhance motion continuity, and motion-speed control to preserve high-fidelity details during rapid movements. Extensive experiments demonstrate that our framework surpasses the conventional duration barrier, generating stable, 720p/30fps videos exceeding one minute with superior temporal stability. Furthermore, the model exhibits robust versatility across five distinct dance genres, conditioned on both audio and textual prompts, establishing a new state-of-the-art in coherent, long-form dance video synthesis.

发表机构

  • Tongyi Lab, Alibaba Group(通义实验室,阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑