通过原子运动生成音乐驱动的舞蹈
Music-to-Dance Generation via Atomic Movements
浏览论文内容
中文总结 AI 辅助
研究音乐驱动的舞蹈生成问题,提出结构感知框架,将编排建模为原子运动序列,经数据分割、聚类及大语言模型处理得到原子运动注释,设计两阶段生成框架,提升舞蹈生成的结构连贯性、节奏对齐和自然度,增强可解释性与可控编辑性。
中文摘要 AI 辅助
音乐驱动的舞蹈生成旨在产生与音乐在节奏上同步且语义上一致的人体运动。近期神经方法虽实现了令人印象深刻的视觉真实感,但将运动建模为连续信号,忽视其组成性质,导致生成的舞蹈结构不连贯且难以控制。本文引入一个结构感知框架,将编排建模为原子运动序列,即作为舞蹈构建块的语义可解释运动事件。通过分割大规模舞蹈数据并聚类成原子运动组,再用大语言模型进行语义重新标记和细化,得到可解释且可复用的原子运动。基于这些注释,设计了一个两阶段生成框架,在原子运动规划阶段预测原子运动的类型、持续时间和时间,在完成阶段生成平滑且风格连贯的运动。实验表明,该方法生成的舞蹈在结构连贯性、节奏对齐和感知自然度方面有显著提升,同时通过显式结构表示增强了可解释性和可控编辑性。
英文摘要
Music-driven dance generation aims to produce human motion that is both rhythmically synchronized and semantically consistent with music. While recent neural approaches have achieved impressive visual realism, they typically model motion as a continuous signal and neglect its compositional nature, making generated dances structurally incoherent and difficult to control. In this work, we introduce a structure-aware framework that models choreography as a sequence of atomic movements-semantically interpretable motion events that serve as the building blocks of dance. To construct this atomic movement vocabulary, we first segment large-scale dance data and cluster them into atomic movement groups. We then employ a large language model to semantically relabel and refine the clusters, yielding a set of interpretable and reusable atomic movements. Based on these atomic movement annotations, we design a two-stage generation framework that mirrors the human choreography process. In the atomic movement planning stage, the model predicts the type, duration, and timing of atomic movements conditioned on the input music, forming a symbolic dance allocation. In the completion stage, a transition-aware generator synthesizes smooth and stylistically coherent motion conditioned on the planned structure. Extensive experiments demonstrate that our method produces dances with significantly improved structural coherence, rhythmic alignment, and perceptual naturalness compared to existing baselines, while providing enhanced interpretability and controllable editing through explicit structural representation.
发表机构
- Wangxuan Institute of Computer Technology, Peking University(北京大学王选计算机研究所)
- School of Electronics Engineering and Computer Science, Peking University(北京大学电子工程与计算机科学学院)
- National Institute of Health Data Science, Peking University(北京大学健康数据科学研究所)
- State Key Laboratory of General Artificial Intelligence, Peking University(北京大学通用人工智能国家重点实验室)
- Beijing Institute for General Artificial Intelligence(北京通用人工智能研究院)
- School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。