Open-UniMo:迈向开放世界中的统一运动-语言理解与生成
Open-UniMo: Towards Unified Motion-Language Understanding and Generation in the Open World
浏览论文内容
中文总结 AI 辅助
提出Open-UniMo统一运动-语言模型,通过扩展词元空间和思维链推理实现双向映射,并借助GRPO优化,在开放世界基准上达到最先进性能,证明生成可促进理解。
中文摘要 AI 辅助
统一运动生成与理解对于能够在开放世界环境中合成和解释人类行为的具身智能系统至关重要。现有的运动-语言模型通常将运动视为语言模型的辅助模态,导致文本主导的表示和有限的跨模态交互。此外,下一个词元预测范式天然不适合长运动序列,因为自回归生成可能会累积预测误差。为了解决这些挑战,我们提出了Open-UniMo,一个在百万级开放世界运动-语言数据上训练的统一大型运动-语言模型(LMLM)。Open-UniMo通过将Qwen约15万文本词元的词汇表扩展6.4万运动词元来促进模态对等,使运动和语言能够共享统一的词元空间。我们进一步引入运动一致的思维链推理作为中间表示,以桥接语言语义和运动动态。Open-UniMo采用两阶段流水线进行训练,其中监督微调建立CoT引导的双向运动-语言映射,组相对策略优化(GRPO)改善语义对齐,同时减轻自回归运动词元生成中的累积误差。为了支持全面评估,我们提出了Open-MoBench,一个VLM引导的基准,用于评估文本到运动(T2M)生成、运动到文本(M2T)理解和双向一致性。大量实验表明,Open-UniMo在传统指标和Open-MoBench上均达到最先进的性能。此外,消融研究揭示,M2T理解主要不受运动词元词汇表大小的限制;相反,将M2T与可学习的T2M生成路径耦合产生更强的跨模态表示,证明在基于AR的运动-语言建模中,生成可以促进理解。
英文摘要
Unified motion generation and understanding is crucial for embodied AI systems that can both synthesize and interpret human actions in open-world environments. Existing motion-language models often treat motion as an auxiliary modality of a language model, leading to text-dominated representations and limited cross-modal interaction. Moreover, the next-token prediction paradigm is not naturally suited to long motion sequences, where autoregressive generation may accumulate prediction errors. To address these challenges, we propose Open-UniMo, a unified Large Motion-Language Model (LMLM) trained on million-scale open-world motion-language data. Open-UniMo promotes modality parity by extending Qwen's vocabulary of about 150K text tokens with 64K motion tokens, enabling motion and language to share a unified token space. We further introduce motion-consistent Chain-of-Thought reasoning as an intermediate representation to bridge language semantics and motion dynamics. Open-UniMo is trained with a two-stage pipeline, where supervised fine-tuning establishes CoT-guided bidirectional motion-language mapping and Group Relative Policy Optimization (GRPO) improves semantic alignment while mitigating cumulative errors in autoregressive motion-token generation. To support comprehensive evaluation, we propose Open-MoBench, a VLM-guided benchmark for assessing text-to-motion (T2M) generation, motion-to-text (M2T) understanding, and bidirectional consistency. Extensive experiments show that Open-UniMo achieves state-of-the-art performance on both conventional metrics and Open-MoBench. Furthermore, ablation studies reveal that M2T understanding is not primarily limited by motion-token vocabulary size; instead, coupling M2T with the learnable T2M generation path yields stronger cross-modal representations, demonstrating that generation can facilitate understanding in AR-based motion-language modeling.
发表机构
- Tsinghua University(清华大学)
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。