arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Token-World:面向机器人操作的视觉-语言模型标记空间世界建模

Token-World: World Modeling in Vision-Language Model Token Space for Robot Manipulation

Chuyao Fu, Xiaowei Chi, Yuhan Rui, Yu-kai Wang, Zezhong Qian, Xiaojie Zhang, Yunfan Lou, Kevin Zhang, Kuangzhi Ge, Chak Wing Mak, Zhiyang Chen, Athena Zhuoming Zhong, Hongyang Chen, Haoran Li, Yike Guo, Sirui Han, Shanghang Zhang

arXiv 2610.00575首次发表:更新:

发表机构

Southern University of Science and Technology; MUKA Robotics; Hong Kong University of Science and Technology; Peking University; University of Pennsylvania(南方科技大学; MUKA Robotics; 香港科技大学; 北京大学; 宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Token-World提出一种在VLM视觉标记压缩空间中建模世界动态的动作条件世界模型,提升机器人操作仿真保真度与策略一致性,并降低延迟。

AI 中文摘要

视觉-语言-动作(VLA)系统的世界模型仿真通常采用预测未来RGB图像并重新编码为策略输入的方式,这会在仿真与下游策略执行之间引入间接接口。我们转而研究是否可以在由VLM视觉标记导出的紧凑、面向策略的状态中建模世界动态。一个关键挑战是原始VLM视觉标记具有高维性,使得高效且准确的自回归动态建模变得困难。为解决这一问题,我们提出了Token-World,一种动作条件世界模型,它将VLM特征压缩为紧凑的标记状态,在该降维空间中学习未来动态,并将预测状态映射回原始面向策略的表示以供下游使用。在多个操作基准测试中,Token-World在开环特征保真度和策略-动作一致性方面优于近期世界模型仿真器,且在长滚动时域内退化更慢。在闭环评估中,其仿真策略性能与参考策略性能的相关性高于Ctrl-World(r=0.794对比0.583),同时所需仿真延迟更低。消融实验进一步表明,紧凑表示设计和维度对未来状态预测有显著影响。代码将在此https URL提供。

英文摘要

A common approach to world-model simulation for vision-language-action (VLA) systems is to predict future RGB observations and then re-encode them into policy inputs, introducing an indirect interface between simulation and downstream policy execution. We instead investigate whether world dynamics can be modeled in a compact, policy-oriented state derived from VLM visual tokens. A key challenge is that raw VLM visual tokens are high-dimensional, making efficient and accurate autoregressive dynamics modeling challenging. To address this, we introduce Token-World, an action-conditioned world model that compresses VLM features into a compact token state, learns future dynamics in this reduced space, and maps predicted states back to the original policy-facing representation for downstream use. Across manipulation benchmarks, Token-World improves open-loop feature fidelity and policy-action consistency over recent world-model simulators, with slower degradation over long rollout horizons. In closed-loop evaluation, its simulated policy performance correlates more strongly with reference policy performance than Ctrl-World ($r=0.794$ vs.\ $0.583$), while requiring lower simulation latency. Ablations further show that compact-representation design and dimensionality substantially affect future-state prediction. Code will be available at https://chuyaofu.github.io/Token-World/.

CommentsSubmitted to IEEE International Conference on Robotics and Automation (ICRA) 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑