arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ARC:机器人基础模型的推理方案

ARC: A Reasoning Recipe for Robot Foundation Models

Gokul Puthumanaillam, Tao Sun, Elie Aljalbout, Moritz Reuss, Zhaoshuo Li, Fabio Ramos, Ankit Goyal, Jenai Xuning Yang

arXiv 2610.12386首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; Stanford University; NVIDIA; University of Sydney; Proception(伊利诺伊大学厄巴纳-香槟分校; 斯坦福大学; 英伟达公司; 悉尼大学; 普洛塞普申公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ARC推理方案,通过推理轨迹、自动标注流水线及模型适配策略,在无需额外机器人演示或大规模训练的情况下,大幅提升了π₀.₅等模型的零样本机器人任务性能,在多个基准上取得显著提升。

AI 中文摘要

当前改进机器人基础模型(RFMs)的主流方法依赖更大规模的模型、更多机器人演示数据以及高成本的大规模训练。本文展示了一种有效且高效的互补方法:合适的推理方案可大幅提升现有最先进RFMs的零样本任务性能,该方案被命名为ARC。它包含三个关键组成部分:推理轨迹、可扩展的自动标注流水线,以及使预训练RFMs利用这些轨迹进行控制的适配策略。首先,研究发现有效的推理轨迹应基于机器人的下一个动作,并解释其因果结构:该动作为何合适、会产生何种效果。其次,这些轨迹可从现有演示中自动生成,无需收集新的机器人数据即可从DROID构建ARC-Trace-DROID。第三,研究展示了最先进的视觉语言动作模型(VLAs)如π₀.₅,以及权重感知模型(WAMs)如Cosmos3-Nano-Policy如何学习利用这些轨迹进行控制,微调与推理过程需适配各模型的架构与能力。使用ARC后,在未增加机器人演示数据或进行基础规模训练的情况下,获得了前所未有的RFMs零样本性能提升。适配后的模型在RoboLab-120和MolmoSpaces上达到新的最先进水平,在RoboLab-Reasoning-50上提升达50个百分点;在真实机器人上,ARC使π₀.₅的任务成功率提升82.2个百分点。项目网站:this https URL

英文摘要

The prevailing approach to improving robot foundation models (RFMs) relies on larger models, more robot demonstrations, and costly training at scale. We show that there exists an effective and efficient complementary approach: the right reasoning recipe can substantially improve the zero-shot task performance of existing state-of-the-art RFMs. We refer to this recipe as ARC. It consists of three key ingredients: a reasoning trace, a scalable automatic labeling pipeline, and a strategy for adapting pretrained RFMs to use these traces for control. First, we find that effective reasoning traces should be grounded in the robot's next action and explain its causal structure: why the action is appropriate and what effect it should produce. Second, we show that these traces can be generated automatically from existing demonstrations, enabling us to construct ARC-Trace-DROID from DROID without collecting new robot data. Third, we show how state-of-the-art VLAs such as $π_{0.5}$ and WAMs such as Cosmos3-Nano-Policy can learn to use these traces for control, with fine-tuning and inference tailored to each model's architecture and capabilities. Using ARC, we obtain gains in zero-shot RFM performance that, to our knowledge, are unprecedented without additional robot demonstrations or foundation-scale training. The adapted models establish a new state of the art on RoboLab-120 and MolmoSpaces, with gains of up to 50 percentage points on RoboLab-Reasoning-50. On real robots, ARC improves $π_{0.5}$'s task success by 82.2 percentage points. Project website: https://arc-robot-reasoning.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑