发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出R^3方法,通过中间训练和基于准则的强化学习将VLM转化为机器人推理器,在两个测试平台上提升了机器人的探索与泛化能力,优于仅指令式的模仿学习基线。
AI 中文摘要
语言推理可让基础模型在测试时将更多计算资源分配给难题,例如需要分解、约束跟踪和未来结果预测的问题。目前尚不清楚该机制能否改善机器人操纵,而长程任务需要跟踪部分进展、推理物体关系、从错误中恢复以及控制嘈杂的低层策略。本文研究能否训练视觉语言模型(VLM)直接以自然语言推理来指导低层操纵策略。我们提出R^3,一种简单的后训练方法,可将现成的VLM转化为机器人推理器:首先用专家生成的推理轨迹对VLM进行中间训练,以初始化所需的推理风格,再基于离线动作数据,通过基于单步骤准则的强化学习来改进推理器。与以往大多使用结构化轨迹作为辅助监督的机器人推理方法不同,R^3训练自由形式的语言推理,以生成测试时的动作指导。我们在Language Table和模拟双手机械臂杂货打包这两个用于研究机器人推理和长程操纵的受控测试平台上实例化R^3。R^3提升了对未见过任务的探索和泛化能力,在两个基准测试中均显著优于仅指令式的模仿学习基线。我们的分析表明,自由形式的语言推理可作为控制低层策略的测试时计算机制。本项目页面可通过该https URL获取。
英文摘要
Reasoning in language allows foundation models to spend more test-time compute on hard problems, such as those requiring decomposition, constraint tracking, and prediction of future consequences. Whether this mechanism can improve robotic manipulation remains unclear, where long-horizon tasks require tracking partial progress, reasoning about object relations, recovering from mistakes, and steering noisy low-level policies. In this paper, we study whether VLMs can be trained to reason directly in natural language to guide low-level manipulation policies. We introduce $R^3$, a simple post-training recipe that turns off-the-shelf VLMs into robotic reasoners: it first mid-trains a VLM on expert-generated reasoning traces to initialize the desired reasoning style, then improves the reasoner with single-step rubric-based RL from offline action data. Unlike prior robotic reasoning methods that mostly use structured traces as auxiliary supervision, $R^3$ trains free-form language reasoning to produce test-time guidance for action. We instantiate $R^3$ on Language Table and simulated bimanual grocery packing, two controlled testbeds for studying robotic reasoning and long-horizon manipulation. $R^3$ improves exploration and generalization across unseen tasks and significantly outperforms instruction-only imitation learning baselines on both benchmarks. Our analyses suggest that free-form language reasoning can function as a test-time compute mechanism for steering low-level policies. Our project page is available at https://robotic-reasoner.github.io/.
Comments42 pages, 23 figures