arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

身体何处保持节拍:结构化动作条件与音乐动态监督用于舞蹈到音乐生成

Where the Body Keeps the Beat: Structured Motion Conditioning and Music Dynamics Supervision for Dance-to-Music Generation

Changchang Sun, Lu Cheng, Yan Yan

arXiv 2610.00726首次发表:更新:

发表机构

University of Illinois Chicago; Penn State University(伊利诺伊大学芝加哥分校; 宾夕法尼亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Dyna2Music潜在流匹配框架,通过结构化动作条件(分解关节速度并分层融合)和潜在动态一致性监督,提升舞蹈到音乐生成的节奏对齐与音频质量。

AI 中文摘要

舞蹈到音乐(D2M)生成旨在合成在时间结构上与给定舞蹈表演对齐的、音乐上合理的配乐。代表性的D2M方法并未明确区分身体部位和频带之间的动作线索,而标准的流匹配缺乏用于监督局部音乐潜在动态的专门目标。为解决这些问题,我们提出了Dyna2Music,一种潜在流匹配框架,结合了结构化动作条件与音乐潜在动态的显式监督。在AIST++上的实证分析量化了空间划分和频率分离如何影响原始音乐到运动学节拍对齐和运动参考密度,为条件设计提供信息。据此,Dyna2Music将关节速度分解为慢速和快速分量,并将由此产生的部位级运动能量与预训练的关节特征分层融合,以条件化音乐生成。为补充这一表示,我们引入了潜在动态一致性(LDC),一个辅助目标,用于匹配单步干净潜在估计与配对参考之间的相邻帧变化幅度。LDC使局部音乐潜在变化成为显式训练目标,而不增加可训练参数或推理计算。Dyna2Music支持可变长度音乐生成,在AIST++和TikTok上的实验表明,与代表性D2M基线相比,节奏对齐和音频质量均有提升。

英文摘要

Dance-to-music (D2M) generation aims to synthesize musically plausible soundtracks whose temporal structure aligns with a given dance performance. Representative D2M methods do not explicitly distinguish motion cues across body parts and frequency bands, while standard flow matching lacks a dedicated objective for supervising local music-latent dynamics. To address these issues, we propose Dyna2Music, a latent flow-matching framework that combines structured motion conditioning with explicit supervision of music-latent dynamics. An empirical analysis on AIST++ quantifies how spatial partitioning and frequency separation affect raw music-to-kinematic beat alignment and motion-reference density, informing the conditioning design. Accordingly, Dyna2Music decomposes joint velocities into slow and fast components and hierarchically fuses the resulting part-wise motion energy with pretrained joint features to condition music generation. To complement this representation, we introduce latent dynamics consistency (LDC), an auxiliary objective that matches adjacent-frame change magnitudes between a single-step clean-latent estimate and the paired reference. LDC makes local music-latent variation an explicit training target without adding trainable parameters or inference computation. Dyna2Music supports variable-length music generation, and experiments on AIST++ and TikTok demonstrate improved rhythmic alignment and audio quality over representative D2M baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑