发表机构
Southern University of Science and Technology; MUKA Robotics; Hong Kong University of Science and Technology; Peking University; University of Pennsylvania(南方科技大学; MUKA Robotics; 香港科技大学; 北京大学; 宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Token-World提出一种在VLM视觉标记压缩空间中建模世界动态的动作条件世界模型,提升机器人操作仿真保真度与策略一致性,并降低延迟。
AI 中文摘要
视觉-语言-动作(VLA)系统的世界模型仿真通常采用预测未来RGB图像并重新编码为策略输入的方式,这会在仿真与下游策略执行之间引入间接接口。我们转而研究是否可以在由VLM视觉标记导出的紧凑、面向策略的状态中建模世界动态。一个关键挑战是原始VLM视觉标记具有高维性,使得高效且准确的自回归动态建模变得困难。为解决这一问题,我们提出了Token-World,一种动作条件世界模型,它将VLM特征压缩为紧凑的标记状态,在该降维空间中学习未来动态,并将预测状态映射回原始面向策略的表示以供下游使用。在多个操作基准测试中,Token-World在开环特征保真度和策略-动作一致性方面优于近期世界模型仿真器,且在长滚动时域内退化更慢。在闭环评估中,其仿真策略性能与参考策略性能的相关性高于Ctrl-World(r=0.794对比0.583),同时所需仿真延迟更低。消融实验进一步表明,紧凑表示设计和维度对未来状态预测有显著影响。代码将在此https URL提供。
英文摘要
A common approach to world-model simulation for vision-language-action (VLA) systems is to predict future RGB observations and then re-encode them into policy inputs, introducing an indirect interface between simulation and downstream policy execution. We instead investigate whether world dynamics can be modeled in a compact, policy-oriented state derived from VLM visual tokens. A key challenge is that raw VLM visual tokens are high-dimensional, making efficient and accurate autoregressive dynamics modeling challenging. To address this, we introduce Token-World, an action-conditioned world model that compresses VLM features into a compact token state, learns future dynamics in this reduced space, and maps predicted states back to the original policy-facing representation for downstream use. Across manipulation benchmarks, Token-World improves open-loop feature fidelity and policy-action consistency over recent world-model simulators, with slower degradation over long rollout horizons. In closed-loop evaluation, its simulated policy performance correlates more strongly with reference policy performance than Ctrl-World ($r=0.794$ vs.\ $0.583$), while requiring lower simulation latency. Ablations further show that compact-representation design and dimensionality substantially affect future-state prediction. Code will be available at https://chuyaofu.github.io/Token-World/.
CommentsSubmitted to IEEE International Conference on Robotics and Automation (ICRA) 2027