发表机构
Sony Computer Science Laboratories; Keio University; Georgia Institute of Technology(索尼计算机科学实验室; 庆应义塾大学; 佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对舞蹈音乐生成中高质量配对数据稀缺问题,提出利用未配对和配对数据的框架,结合预训练单峰编码器、对比预训练及控制网络模块,实验证明该方法提升了舞蹈音乐对齐和音频质量,优于现有技术。
AI 中文摘要
舞蹈音乐生成对编排支持和自动伴奏等应用很有前景,身体动作与声音的时间协调至关重要。使用人体关节位置作为运动表示很有吸引力。但高质量配对舞蹈音乐数据稀缺,难以仅从配对数据训练端到端模型。我们提出一个舞蹈条件音乐生成框架,利用未配对和配对数据,结合预训练单峰编码器、节拍引导对比预训练及控制网络风格的条件模块。在AIST++上的实验表明该技术改善了舞蹈音乐对齐和音频质量,优于现有方法。
英文摘要
Dance-to-music generation is a promising task for applications such as choreography support and automatic accompaniment, where temporal coordination between body movement and sound is essential. In particular, using human joint positions as the motion representation is attractive because they explicitly capture body dynamics while being lightweight, privacy-preserving, and easy to integrate with motion capture and pose-estimation pipelines. A central challenge in this setting, however, is the scarcity of high-quality paired dance-music data, since collecting accurately synchronized pairs is costly and often constrained by copyright and performance rights. This makes it difficult to train end-to-end models solely from paired data. To address this issue, we propose a dance-conditioned music generation framework that efficiently exploits both unpaired and paired data. Our method combines pretrained unimodal encoders for motion and music, beat-guided contrastive pretraining to align their feature spaces, and a ControlNet-style conditioning module on top of a pretrained text-to-audio diffusion model. Experiments on AIST++ demonstrate that the proposed techniques improve both dance-music alignment and audio quality, as confirmed by quantitative and qualitative evaluations. Compared to a state-of-the-art method, our approach achieves superior dance alignment performance and competitive audio quality. Code is available at https://github.com/kmraven/AudioLDM-ControlNet .
Comments7 pages, 1 figure