发表机构
Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MotionMaestro提出掩码运动分词统一框架,通过观测模式统一多任务条件,采用三阶段训练和观测损失,在大型数据集上实现多种运动生成任务的最先进性能。
AI 中文摘要
人体运动生成在角色动画、虚拟环境和具身交互等应用中扮演着重要角色。尽管现有方法已取得显著进展,但其中许多是针对单个任务开发的,包括文本到运动、姿态条件生成和轨迹控制。虽然这些任务涉及不同类型的条件,但一个能够在共同表示内处理它们的统一框架将大大简化运动生成系统。我们观察到,多样的运动条件可以自然地表述为运动序列上的不同观测模式,其中每个任务对应一种特定的掩码策略。基于这一见解,我们引入了MotionMaestro,一个统一的运动生成框架,通过掩码运动分词学习完整运动和异构部分观测的共享表示。MotionMaestro采用三阶段训练策略:首先学习掩码运动分词器,然后在干净运动上优化其重建能力,最后在学得的潜在空间中训练条件流匹配生成器。此外,我们引入了观测图和观测损失,以在生成过程中明确保留提供的运动条件。凭借这种统一表示和条件机制,MotionMaestro支持文本引导和无条件合成、姿态条件和部分补全、时间插值、轨迹控制和运动延续。在大型RoMo和MotionMillion数据集上的实验表明,其在多种运动生成任务中达到了最先进的性能。
英文摘要
Human motion generation plays an important role in applications such as character animation, virtual environments, and embodied interaction. While existing approaches have achieved remarkable progress, many of them are developed for individual tasks, including text-to-motion, pose-conditioned generation, and trajectory control. Although these tasks involve different types of conditions, a unified framework capable of handling them within a common representation would greatly simplify motion generation systems. We observe that diverse motion conditions can be naturally formulated as different observation patterns over motion sequences, where each task corresponds to a specific masking strategy. Based on this insight, we introduce MotionMaestro, a unified motion generation framework that learns a shared representation for complete motions and heterogeneous partial observations through masked motion tokenization. MotionMaestro employs a three-stage training strategy that first learns a masked motion tokenizer, then refines its reconstruction ability on clean motions, and finally trains a conditional flow-matching generator in the learned latent space. Furthermore, we introduce an observation map and an observation loss to explicitly preserve provided motion conditions during generation. With this unified representation and conditioning mechanism, MotionMaestro supports text-guided and unconditional synthesis, pose conditioning and partial completion, temporal interpolation, trajectory control, and motion continuation. Experiments on the large-scale RoMo and MotionMillion datasets show state-of-the-art performance across diverse motion generation tasks.
CommentsPlease visit our project page https://kaist-viclab.github.io/MotionMaestro_site/