arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15984cs.CV

用于真实世界运动语言模型的即插即用二维运动接口

A Plug-and-Play 2D Motion Interface for Real-World Motion Language Models

  • Toyota Technological Institute(丰田工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Kaname Yokoyama, Norimichi Ukita

AI总结:

针对MoLMs依赖三维运动输入、真实世界适用性受限的问题,提出即插即用二维运动接口,经实验验证其性能接近三维输入,优于从零训练的二维输入方案,为MoLMs部署提供实用途径。

AI中文摘要:

运动语言模型(MoLMs)通常通过对三维运动进行分词,再用语言模型处理得到的分词来理解人类运动。然而,从单目视频中获取准确的三维运动颇具挑战性,限制了其在真实世界中的适用性。为解决该问题,我们提出一种即插即用的二维运动接口,使经过三维预训练的MoLMs能够接受二维运动输入,且无需修改或微调原始模型。在公开数据集上开展的实验表明,我们的方法在多个MoLMs上取得了与三维运动输入相当的性能,且优于从零开始在二维运动上训练MoLMs的效果。我们还构建了一个单目真实世界视频运动评估数据集,并引入了真实视频适配器,证明在评估的单目姿态估计设置下,二维运动相比三维运动更具实用性。这些结果表明,二维运动为将MoLMs部署到真实世界运动理解场景提供了实用接口。代码可在该https URL获取。

英文摘要:

Motion Language Models (MoLMs) typically understand human motions by tokenizing 3D motion and processing the resulting tokens using a language model. However, obtaining accurate 3D motions from monocular videos is challenging, limiting their real-world applicability. To address this issue, we introduce a plug-and-play 2D Motion Interface that enables 3D-pretrained MoLMs to accept 2D motion inputs without modifying or fine-tuning the original models. Experiments on public datasets show that our method achieves performance comparable to 3D motion inputs across multiple MoLMs and outperforms training MoLMs from scratch on 2D motions. We further construct a monocular real-world video motion evaluation dataset and introduce a real-video adapter, demonstrating the usefulness of 2D motions over 3D motions under the evaluated monocular pose-estimation setting. These results suggest that 2D motion provides a practical interface for deploying MoLMs in real-world motion understanding settings. Code is available at https://github.com/irajisamurai/2D-Motion-Interface.

补充信息

↑