arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13118cs.AIcs.SD

CMA-OT: 用于舞蹈到音乐生成的层级专家监督

CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation

Jinting Wang, Chenxing Li, Dong Yu, Li Liu

首次发表
浏览论文内容

中文总结 AI 辅助

针对舞蹈到音乐生成中稀疏线索与密集音乐信息语义不匹配的问题,提出CMA-OT方法,利用外部音乐专家进行层级监督,结合课程引导多尺度学习和尺度感知最优传输对齐,在两个数据集上达到最先进性能。

中文摘要 AI 辅助

舞蹈到音乐(D2M)生成旨在合成与舞蹈视频在节奏和风格上对齐的音乐。一个关键挑战源于稀疏的舞蹈线索(如节奏和风格)与音乐创作所需的密集信息(包括结构、配器和表现力动态)之间的语义不匹配。现有方法通常依赖这些稀疏线索,并且仅对最终音频输出进行监督,导致音乐表示学习不佳,生成的音乐音乐性和结构连贯性有限。为解决这些问题,我们提出了课程引导的多尺度表示对齐与尺度感知最优传输(CMA-OT),这是一种新颖的范式,利用外部音乐专家对生成器的潜在特征提供层级监督,弥合语义差距并增强表示学习。为了有效整合层级监督,我们引入了一种课程引导的多尺度学习策略,逐步将音乐知识从专家转移到音乐生成器,实现稳定且有效的表示学习。此外,为了适应不同专家尺度之间的语义和结构变化,并在时间不匹配下实现细粒度对齐,我们提出了一种尺度感知的最优传输对齐机制,该机制建模层级专家表示与生成器潜在特征之间的软对应关系。在两个数据集上的大量实验表明,CMA-OT在节奏同步、感知质量和整体音乐生成方面达到了最先进的性能。

英文摘要

Dance-to-music (D2M) generation aims to synthesize music that is rhythmically and stylistically aligned with dance videos. A key challenge arises from the semantic mismatch between sparse dance cues, such as rhythm and style, and the dense information required for music composition, including structure, instrumentation, and expressive dynamics. Existing methods typically rely on these sparse cues and supervise only the final audio output, resulting in poorly learned music representations and generated music with limited musicality and structural coherence. To address these issues, we propose Curriculum-guided Multi-scale representation Alignment with scale-aware Optimal Transport (CMA-OT), a novel paradigm that leverages an external music expert to provide hierarchical supervision for the generator's latent features, bridging the semantic gap and enhancing representation learning. To effectively incorporate hierarchical supervision, we introduce a curriculum-guided multi-scale learning strategy that progressively transfers musical knowledge from the expert to the music generator, enabling stable and effective representation learning. Moreover, to accommodate the semantic and structural variations across different expert scales and achieve fine-grained alignment under temporal mismatch, we propose a scale-aware optimal transport alignment mechanism, which models soft correspondences between hierarchical expert representations and the generator's latent features. Extensive experiments on two datasets demonstrate that CMA-OT achieves state-of-the-art performance in rhythmic synchronization, perceptual quality, and overall music generation.

发表机构

  • HKUST(GZ)(香港科技大学(广州))
  • Tencent(腾讯)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑