arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27581cs.LG

MUGEN:用于高效运动理解与生成的统一框架

MUGEN: A Unified Framework for Efficient Motion Understanding and Generation

Zhankai Ye, Yukai Jin, Bingyang Wei, Bofan Li, Yusen Wu, Fangyi Li, Shangqian Gao, Xin Liu

首次发表
浏览论文内容

中文总结 AI 辅助

MUGEN是一个无码本的统一运动-语言框架,通过自适应长度自编码器实现高效运动压缩,在HumanML3D和SnapMoGen数据集上的生成、检索与对齐指标均表现优异,兼顾效率与性能。

中文摘要 AI 辅助

将人类运动与语言建立对应关系,以及将语言与运动建立对应关系,是构建能够理解、生成和交流人类行为的物理AI系统的核心步骤。统一的运动-语言系统最初通过共享的离散运动码本耦合这两个方向,但量化过程限制了生成质量。最强的生成器以不断增长的成本换取质量:堆叠的残差码本扩大了表示;掩码解码阶段、长自回归展开以及数十到数百步的去噪链延长了推理过程;其中的连续潜变量设计也仅通过迭代扩散头才能得到潜变量,且所有这些解码机制都无法服务于理解。因此,我们提出MUGEN,这是一个统一的运动-语言框架,无需付出上述任何成本:没有码本,仅一次采样。单个自适应长度的自编码器将任意长度的运动压缩为少量连续潜变量槽,这是系统唯一的运动表示:语言模型为文本到运动任务生成这些槽,为运动理解任务读取这些槽。深度路由的隐藏状态让每个槽从所需的Transformer深度读取信息,且校准后的头预测所有潜变量的联合分布,因此一次采样即可承载描述所允许的文本条件跨槽变异。在K个语言模型步、一次采样和一次解码器传递的解码成本下,MUGEN在HumanML3D上的FID指标优于语言模型基线,同时在标准评估器下将检索精度提升至超过真实运动参考值,取得了最佳的CIDEr和BLEU@4分数,并在SnapMoGen的所有检索和对齐指标上超越了离散令牌的现有技术水平。

英文摘要

Grounding human motion in language, and language in motion, is a central step toward physical AI systems that can understand, generate, and communicate human behavior. Unified motion--language systems first coupled the two directions through a shared discrete motion codebook, but quantization limits generation quality. The strongest generators buy quality back at growing cost: stacked residual codebooks enlarge the representation; masked decoding stages, long autoregressive rollouts, and denoising chains of tens to hundreds of steps stretch inference; even the continuous-latent designs among them reach their latent only through an iterative diffusion head; and none of this decoding machinery serves understanding. We therefore propose MUGEN, a unified motion--language framework that pays neither cost: no codebook, one draw. A single adaptive-length autoencoder compresses any-length motion into a few continuous latent slots, the system's only motion representation: the language model generates them for text-to-motion and reads them back for motion understanding. Depth-routed hidden states let each slot read from the transformer depth it needs, and a calibrated head predicts a joint distribution over the full latent set, so a single draw carries the text-conditional, cross-slot variation a description permits. At a decoding cost of K language-model steps, one draw, and one decoder pass, MUGEN leads language-model baselines on FID on HumanML3D while raising retrieval precision above the real-motion reference under the standard evaluator, achieves the best CIDEr and BLEU@4 scores, and surpasses the discrete-token state of the art on every retrieval and alignment metric on SnapMoGen.

发表机构

  • Florida State University(佛罗里达州立大学)
  • Texas Christian University(得克萨斯基督教大学)
  • University of Miami(迈阿密大学)
  • University of Pennsylvania(宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

↑