MotionCanvas:从可组合运动学线索中学习隐式运动规划
AnimateCanvas: Learning Implicit Motion Planning from Composable Kinematic Cues
浏览论文内容
中文总结 AI 辅助
MotionCanvas提出一种基于共享运动画布和流匹配模型的隐式运动规划方法,通过组合线索采样器实现多类型运动学线索的连贯融合,在多项任务上达到最先进水平。
中文摘要 AI 辅助
专业角色动画既需要自然的运动,也需要精确且多功能的控制。例如,创作者通常需要定义特定动作的时间、控制角色手臂摆动的运动范围以及角色行走的路线,就像在“运动画布”上指定各种运动学运动线索一样。这促使我们提出MotionCanvas,一个支持线索条件隐式运动规划的模型,能够忠实且连贯地将所有线索(密集或稀疏、完整或部分)连接成一个全身运动序列。具体而言,MotionCanvas在共享的运动画布上表示异构的运动学线索,其中位置和旋转值在身体关节和时间上被指定。一个共享的流匹配模型生成以该画布为条件的运动,并可选地结合语言和输入运动;线索插补在训练和采样过程中保持指定的画布值不变。为了学习不同线索集之间的连贯补全,我们使用一个组合线索采样器进行训练,该采样器变化线索应用的时间、指定的位置或旋转以及它们的组合方式。这些设计共同使一个单一生成器能够合成全局连贯的动作,同时满足兼容的异构线索。我们通过时间、根部和身体部位线索(单独及组合)以及语言引导编辑来测试这种规划能力。我们自然地将此评估扩展到序列生成和运动修复,因为两者都需要从运动学线索组织连贯运动的相同能力。在这些评估中,MotionCanvas在受控运动质量、混合线索遵循、序列生成、指令编辑和运动修复方面均取得了最先进的结果,同时保持了其文本到运动的能力。
英文摘要
Professional character animation requires both natural motion and precise, versatile control. For example, creators often define the timing of a specified action, control the motion range of the character's arm swing, or specify the route the character walks through--effectively placing various kinematic cues on a motion canvas. This motivates us to propose AnimateCanvas, a model that supports cue-conditioned implicit motion planning to faithfully and coherently connect all cues, dense or sparse, full or partial, into one full-body motion sequence. Specifically, AnimateCanvas represents heterogeneous kinematic cues on a shared motion canvas, where position and rotation values are specified across body joints and time. A shared flow-matching model generates motion conditioned on this canvas, with optional language and input motion; cue imputation keeps the specified canvas values fixed in both training and sampling. To learn coherent completion across different cue sets, we train with a compositional cue sampler that varies the timing of cue application, the positions or rotations specified, and how they are combined. Together, these designs enable a single generator to integrate heterogeneous kinematic cues into coherent full-body actions, giving creators fine-grained control over selected frames, joints, and position or rotation channels. We evaluate this planning ability on temporal, root, and body-part cues--alone and combined--as well as language-guided editing, and naturally extend it to sequential generation and motion repair. AnimateCanvas achieves state-of-the-art results in temporal completion, spatial control, sequential generation, language-guided editing, and motion repair, while retaining strong text-to-motion capability.
发表机构
- Zhejiang University(浙江大学)
- Tencent(腾讯)
- Peking University(北京大学)
- Sun Yat-sen University(中山大学)
- Zhejiang Lab(浙江省实验室)
机构由 AI 辅助整理,请以论文原文为准。