arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于训练时适配与推理时引导的生成式运动模型可控夸张

Controllable Exaggeration for Generative Motion Models via Training-Time Adaptation and Inference-Time Guidance

Amirhossein Zamani, Arianna Rampini, Bruno Roy

arXiv 2610.12316首次发表:更新:

发表机构

Autodesk Research; Mila – Quebec AI Institute; Concordia University(奥多比研究机构; 米拉-魁北克人工智能研究所; 康考迪亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对现有运动生成模型忽略动画夸张原则的问题,提出训练时微调与推理时引导的双阶段框架,可在保留运动合理性的同时生成更具表现力的夸张运动。

AI 中文摘要

近期的运动生成模型已展现出合成物理上合理的角色运动的强大能力,但往往忽略了专业动画师用于构建和设计动画作品的成熟动画原则。理解并将这些原则融入运动生成流程,对于生成既适用于物理基础应用、又能满足角色动画社区需求的运动至关重要,这能让角色不仅以物理合理的方式运动,还显得鲜活、富有表现力且引人入胜。为缩小这一差距,我们聚焦于动画的夸张原则,研究如何将其融入现代运动生成流程以生成更具表现力的角色运动。为此,我们引入一个在现有运动生成流程两个阶段运行的框架:第一阶段在训练时引入夸张,我们对预训练的文本到运动模型在我们整理的夸张数据集上执行监督微调;第二阶段在推理时运行,我们:(i)基于动态运动原语(DMPs)提出夸张的数学表述;(ii)利用该表述作为夸张引导信号,引导现有的扩散模型和流匹配文本到运动生成模型生成夸张运动,无需额外训练。通过与三个强大的运动生成模型进行定性和定量评估,我们表明,我们的方法在保留中性参考运动意图和物理合理性的同时,生成了更夸张、更具表现力的运动。

英文摘要

Recent motion generative models have demonstrated strong capabilities in synthesizing physically plausible character motion, but often overlook established animation principles used by professional animators to ground and design their animation work. Understanding and incorporating these principles into motion generative pipelines is essential for producing motions that serve not only physically grounded applications but also the needs of the character animation community. This enables the creation of characters that not only move in physically plausible ways but also feel alive, expressive, and engaging. To close this gap, we focus on the Exaggeration principle of animation and investigate how it can be incorporated into modern motion generative pipelines to produce more expressive character motions. To this end, we introduce a framework that operates at two stages of existing motion generative pipelines. The first stage introduces exaggeration during training, where we perform supervised fine-tuning of pre-trained text-to-motion models on our curated exaggeration dataset. The second stage operates at inference time, where we: (i) introduce a mathematical formulation of exaggeration based on dynamic movement primitives (DMPs); and (ii) leverage this formulation as an exaggeration guidance signal to guide existing diffusion and flow-matching text-to-motion generation models toward exaggerated motion without additional training. Through qualitative and quantitative evaluations against three strong motion generation models, we show that our methods generate more exaggerated and expressive motions while preserving neutral reference motion intent and physical plausibility.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑