arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WorldToken:面向机器人模仿学习的时间优先序列建模

WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning

Chunkai Yang, Andong Yang, Di Huang, Chao Gao, Guyue Zhou

arXiv 2608.22591首次发表:更新:

发表机构

School of Remote Sensing and Information Engineering, Wuhan University; Tsinghua University; Institute for AI Industry Research, Tsinghua University(武汉大学遥感信息工程学院; 清华大学; 清华大学人工智能产业研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出时间优先序列建模方法WorldToken,融合多类观测生成世界token序列,结合因果时间Transformer与扩散动作头,在机器人模仿学习任务上验证了其可行性及数据缩放、时间上下文特性。

AI 中文摘要

机器人策略在每个决策步骤都会接收到异构观测值,但序列模型在时间维度上组织这些输入的方式存在差异。我们提出WorldToken,一种时间优先的策略实例化方法,将每个策略时间步内的多视角图像、本体感知信息和任务条件融合为一个世界token。因果时间Transformer对生成的世界token序列进行建模,扩散动作头生成动作块。在23个RoboCasa任务上,除冻结的预训练CLIP文本编码器外,从头训练的8530万参数策略,利用每个任务2900个生成的演示,实现了59.45%的平均闭环成功率。对5种数据集规模、5种模型规模和2个训练随机种子进行的完整因子扫描显示,额外目标域数据会带来持续增益,而超过中等模型规模后会出现收益递减。在相同检查点历史截断条件下,将可见历史减少至1或2个策略时间步会降低所有50个RoboCasa策略的闭环成功率。在RMBench Blocks Ranking任务中,将可见历史从146减少至8秒,评估器成功率从95%降至28%,而探索性扩展滚动能维持参考交换序列超过850秒。这些结果确立了完整WorldToken实例化的经验可行性,并表征了其在测试方案下的数据缩放和时间上下文行为,但未确立其相对于其他序列组织方式的优越性,也未分离完整实现中驱动观测性能的组件。

英文摘要

Vision-language-action policies inherit both capabilities and input representations from pretrained vision-language models, showing great potential for robotic manipulation across diverse industrial settings. As these policies increasingly use interaction history, organizing the representation of historical observations determines how experiences enter temporal context and how relationships across time are modeled, which is a generally ignored challenge in previous works. In this work, we believe solving this challenge requires an architectural reconstruction and propose WorldToken. WorldToken encodes each timestep's observations into one world token, processes the resulting history with a causal Transformer, and generates action chunks with a diffusion action head. Unified token enables long horizon tasks while relieving the memory requirement of the temporal backbone, hence improving performance. This design also offers high interpretability and allows advances in language modeling, such as pre-training and scaling, to be transferred to robot interaction policies. In RoboCasa experiments, WorldToken successfully handles most tasks with 85M parameters and achieves 59.4% mean closed-loop success close to $π_{0.5}$ with 3.35B parameters. On the memory benchmark RMBench Blocks Ranking, WorldToken can reach the context of two minutes and achieves a success rate of 95%. In addition, holdout action RMSE is well described by power-law fits, and its closed-loop success rate improves consistently with increasing training data size in a study involving approximately 350,000 closed-loop evaluation episodes across 50 trained policies on RoboCasa, showing its scaling potential. We also conduct experiments to analyze information preservation and history use in WorldToken, providing empirical grounding for future work.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑