arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16885cs.RO

$τ_0$-VLA:一种具有世界模型引导测试时计算能力的分层机器人基础模型

$τ_0$-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation

  • Shanghai Innovation Institute(上海创新研究院)
  • Agibot Finch(智元 Finch(智元鹦鹉))
  • The Chinese University of Hong Kong(香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

Xiaowei Cai, Yunuo Cai, Bingao Chen, Jingxiao Chen, Zhi Chen, Siyuan Feng, Tengyu Hou, Jingshun Huang, Han Jiang, Runkun Ju, Dong Li, Mingxiang Li, Shaowei Li, … 展开作者

Xiaowei Cai, Yunuo Cai, Bingao Chen, Jingxiao Chen, Zhi Chen, Siyuan Feng, Tengyu Hou, Jingshun Huang, Han Jiang, Runkun Ju, Dong Li, Mingxiang Li, Shaowei Li, Xinchen Li, Yifan Li, Yi Liu, Zhongyuan Liu, Jianlan Luo, Junwen Miao, Ruiqi Ni, Buqing Nie, Mingjie Pan, Xinlin Ren, Jianheng Song, Jiaxu Wang, Peiqi Wang, Sen Wang, Xiaoyan Wang, Dafeng Wei, Dongming Wu, Pengwei Xie, Pu Yang, Hangjian Ye, Xiangyu Yue, Jinyu Zhang, Qinglin Zhang, Xueyong Zhao, Pengfei Zhou, Yue Zhou

AI总结:

$τ_0$-VLA是一种分层机器人基础模型,通过世界模型引导的测试时计算分配额外资源优化子任务生成,经40115小时异构真实数据训练后,可提升长时程机器人操作的闭环成功率。

AI中文摘要:

长 horizon(长时程)机器人操作要求机器人既能可靠执行单个技能,又能在扩展任务中连贯地将这些技能排序。大多数分层视觉-语言-动作(VLA)模型通过单次前向传递做出每个此类决策,没有机制为困难或重要的选择分配额外计算资源。我们提出$τ_0$-VLA,一种分层机器人基础模型,它将高层子任务生成为由世界模型引导的测试时计算的可扩展推理问题。在每个推理步骤,高层策略利用执行记忆生成子任务,必要时会在多个备选方案中搜索后再确定输出。低层策略随后在多个机器人 embodiment(实体)上执行生成的子任务。该策略在包含40115小时异构真实世界数据的多模态协同训练数据上进行训练。在域内和分布偏移设置中,分配额外测试时计算显著提升了下一个子任务的预测准确率,且这些提升转化为长时程机器人操作任务更高的闭环成功率。

英文摘要:

Long-horizon robot manipulation requires a robot to both execute individual skills reliably and sequence them coherently over extended tasks. Most hierarchical vision-language-action (VLA) models make each such decision with a single forward pass, leaving no mechanism to allocate additional computation to difficult or consequential choices. We introduce $τ_0$-VLA, a hierarchical robot foundation model that formulates high-level subtask generation as a compute-scalable inference problem through world-model-guided test-time computation. At each inference step, the high-level policy uses execution memory to generate a subtask and, when needed, searches over alternatives before committing to its output. A low-level policy then executes the generated subtask across multiple robot embodiments. The policy is trained on 40,115 hours of heterogeneous real-world data with multimodal co-training. Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.

补充信息

相关深度报道

↑