arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24334cs.CVcs.CLcs.GR

SeMoCo:面向运动语言建模的语义优先运动编解码器

SeMoCo: A Semantic-First Motion Codec for Motion Language Modeling

Tianlv Huang, Hetian Guo, Ziyi Cai, Song Wang, Yanping Zhang, Zipei Fan, Xuan Song, Guangming Wu, Xin Zheng

首次发表
浏览论文内容

中文总结 AI 辅助

SeMoCo是面向运动语言建模的语义优先运动编解码器,搭配双轴运动生成器,构建了Ω-MotionVerse数据集,在重构精度与文本到运动生成任务中表现优异。

中文摘要 AI 辅助

离散运动表示已显著推动自回归文本到运动生成的发展,但大多数运动分词器针对重构优化,未按语义角色明确分配容量,导致动作级语义与细粒度运动学细节需通过同一重构驱动层级编码。本文提出SeMoCo语义优先运动编解码器,及用于语言条件运动生成的双轴运动生成器,每个运动令牌包含1个语义令牌与残差运动学令牌序列,生成器对跨时间的语义进展建模并自回归优化残差项。同时构建基于SOMA表示统一的大规模多源人体运动数据集Ω-MotionVerse。对比实验显示,SeMoCo在被测编解码器中重构精度最优,且文本到运动生成的良好结果验证了其运动令牌对下游生成任务的有效性。

英文摘要

Discrete motion representations have substantially advanced autoregressive text-to-motion generation. However, most motion tokenizers are optimized for reconstruction and do not explicitly allocate capacity according to semantic role. Action-level meaning and fine-grained kinematic detail must therefore be encoded through the same reconstruction-driven hierarchy. We introduce SeMoCo, a semantic-first motion codec, together with a dual-axis motion generator for language-conditioned motion generation. Each motion token contains one semantic token and a residual sequence of kinematic tokens. The generator models semantic progression across time and autoregressively refines the residual entries. We also construct $Ω$-MotionVerse, a large-scale, multi-source human-motion dataset unified under the SOMA representation. Across the reported comparisons, SeMoCo achieves the best reconstruction accuracy among the compared codecs, while strong text-to-motion results demonstrate the effectiveness of its motion tokens for downstream generation.

发表机构

  • Jilin University(吉林大学)
  • Frontier Robotics(前沿机器人公司)
  • Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))

机构由 AI 辅助整理,请以论文原文为准。

↑