发表机构
College of Software, Jilin University; College of Computer Science and Technology, Zhejiang University; ReLER, AAII, University of Technology Sydney; College of Computer Science and Technology, Jilin University; School of Computer Science and Technology, Zhejiang Normal University; School of Information and Communication Engineering, Dalian University of Technology; School of Biomedical Engineering, Shenzhen University Medical School, Shenzhen University; School of Computer Science and Informatics, Cardiff University(吉林大学软件学院; 浙江大学计算机科学与技术学院; 悉尼科技大学AAII ReLER实验室; 吉林大学计算机科学与技术学院; 浙江师范大学计算机科学与技术学院; 大连理工大学信息与通信工程学院; 深圳大学医学部生物医学工程学院; 卡迪夫大学计算机科学与信息学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
BeatDance提出一种基于扩散的层次化时空建模框架,通过解耦注意力与循环一致性学习,实现音乐节拍与三维舞蹈动作的精确对齐,并在两个基准数据集上超越现有方法。
AI 中文摘要
从音乐生成逼真的三维舞蹈是一项具有挑战性的任务,它要求在与音乐节奏精确同步的同时,捕捉人体运动的时空复杂性。尽管现有方法能够生成物理上合理的舞蹈动作,但它们往往难以实现与音乐(如节拍)的精确对齐。为解决这一局限,我们提出了一种新颖的基于扩散的框架BeatDance,该框架包含两个组成部分:1)我们提出了一种层次化解耦注意力(HDA)模块,该模块首先将人体姿态与时间动态的学习进行解耦,然后采用层次化结构来捕捉短期与长期依赖关系,从而增强时空建模能力。2)我们通过引入一个辅助的舞蹈到音乐模块,采用循环一致性学习。在训练过程中,重建音乐与原始音乐之间的差异会产生更强的损失信号,有效促进音乐与舞蹈动作之间的一致性。大量实验结果表明,我们提出的方法在两个基准数据集上优于近期具有竞争力的方法。
英文摘要
Generating realistic 3D dance from music is a challenging task that requires accurate synchronization with musical rhythms while capturing the spatial complexity of human motion. Although existing methods can generate physically plausible dance motions, they often struggle to achieve precise alignment with music, such as the beat. To address this limitation, we propose a novel diffusion-based framework, BeatDance, with two components: 1) We present a Hierarchical Decoupled Attention (HDA) module, which first disentangles the learning of human pose and temporal dynamics. A hierarchical structure is then employed to capture both short-term and long-term dependencies, thereby enhancing spatial-temporal modeling. 2) We adopt cycle-consistent learning by introducing an auxiliary dance-to-music module. During training, discrepancies between the reconstructed and original music induce a stronger loss signal, effectively encouraging the consistency property between the music and dance motion. Extensive experimental results demonstrate that our proposed approach outperforms recent competitive methods on two benchmark datasets.
CommentsPublished in Pattern Recognition
Journal refPattern Recognition, Volume 180, 114344 (2026)
DOI:10.1016/j.patcog.2026.114344