发表机构
DisneyResearch|Studios; The Blavatnik School of Computer Science and AI, Tel-Aviv University; ETH Zürich(迪士尼研究工作室; 特拉维夫大学布拉瓦尼克计算机科学与人工智能学院; 苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出带双目标损失函数的生成式扩散框架,结合自有数据集与新指标,实现从音频生成厘米级精度的逼真打鼓动作,泛化能力优于现有方法。
AI 中文摘要
音乐驱动的角色动画可赋能并拓展娱乐与互动教育领域的变革性应用。然而,从音频合成逼真的打鼓动作仍具挑战性,因为高加速度动态特性与极端时空精度需求之间存在固有矛盾。现有方法通常依赖动作匹配或MIDI输入,难以泛化到多样的真实世界音频。此外,该领域缺乏能够区分精确打鼓与噪声动作的标准化评估指标。在本文中,我们引入一种生成式扩散框架,该框架具有双目标损失函数,可将骨骼完整性与鼓槌精度解耦,从而在不牺牲自然身体动态的情况下实现厘米级的鼓槌精度。此外,利用我们自己的数据集和数据增强策略,该模型可泛化到非精选的野外音频。为严格评估性能,我们提出两个新指标:用于量化空间精度的击打-目标距离,以及用于评估时间对齐的音频-动作相关分数。我们的定量分析和用户研究表明,我们的系统生成的高质量动作通常与真实表演难以区分。
英文摘要
Music-driven character animation enables and enhances transformative applications in entertainment and interactive education. However, synthesizing realistic drumming motion from audio remains challenging due to the inherent tension between high-acceleration dynamics and the need for extreme spatial-temporal precision. Existing approaches, often reliant on motion matching or MIDI input, struggle with generalizing to diverse real-world audio. Moreover, the field lacks standardized evaluation metrics capable of distinguishing precise drumming from noisy motion. In this paper, we introduce a generative diffusion framework featuring a dual-objective loss function that decouples skeletal integrity from drumstick precision, thus enabling centimeter-level stick precision without sacrificing natural body dynamics. Additionally, leveraging our own dataset and data augmentation strategy, the model generalizes to non-curated, in-the-wild audio. To rigorously evaluate performance, we propose two novel metrics: an impact-to-target distance to quantify spatial precision and an audio-motion correlation score to assess temporal alignment. Our quantitative analysis and user studies demonstrate that our system generates high-quality motion that is often indistinguishable from ground-truth performances.
CommentsBest Paper Award at the 25th ACM SIGGRAPH / Eurographics Symposium on Computer Animation (SCA 2026). For Supplementary Video, see https://studios.disneyresearch.com/2026/08/18/generalized-audio-driven-synthesis-of-precise-drummer-motion/