arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

寥寥数语,成效显著:语言引导的机器人策略合成

A Few Words Go a Long Way: Language Guided Robot Policy Synthesis

Daphne Chen, Archit Ritesh Jain, Eric Goossen, Emma Romig, Michael Murray, Nick Walker, Maya Cakmak

arXiv 2607.23784首次发表:更新:

发表机构

University of Washington; Microsoft Research; Massachusetts Institute of Technology(华盛顿大学; 微软研究院; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对视觉-语言-动作模型难以解释等问题,提出ARCHITECT框架,利用大语言模型编码代理合成模块化机器人程序,通过人类自然语言修正引导策略并积累技能库,在机器人基准评估中表现出色,提供了新的机器人学习方案。

AI 中文摘要

虽然视觉-语言-动作模型展现出令人印象深刻的零样本操作能力,但它们本质上仍是难以解释、适应或修正的黑箱策略。本文提出ARCHITECT框架,将机器人策略获取视为交互式程序合成任务。利用大语言模型编码代理的推理能力合成模块化机器人程序,通过迭代过程让人类监督者用自然语言修正引导策略,修正基于执行轨迹并融入持久技能库。在Franka Panda机器人基准评估中,ARCHITECT在复杂长时任务上优于现有模型和程序合成基线,展示了可转向且数据高效的黑箱机器人学习替代方案。

英文摘要

While vision-language-action models have demonstrated impressive zero-shot manipulation capabilities, they remain fundamentally black box policies that are difficult to interpret, adapt, or correct when they inevitably fail. In this work, we propose ARCHITECT, a framework that treats robot policy acquisition as an interactive program synthesis task. ARCHITECT leverages the reasoning capabilities of LLM coding agents to synthesize modular robot programs that utilize a suite of perception and control tools. Unlike end-to-end models where distribution shift leads to unpredictable, cascading failures, our modular architecture allows users to isolate failures and localize feedback at the level of abstraction required. We introduce an iterative process where a human supervisor provides natural language corrections to steer the policy. These corrections are grounded in the policy code by program execution traces and distilled into a persistent skill library, a form of long-term in-context learning which enables the agent to accumulate a repertoire of reusable, interpretable behaviors. In a benchmark evaluation on a Franka Panda robot, ARCHITECT outperforms state-of-the-art VLA models and program synthesis baselines on complex, long-horizon tasks, including articulated object manipulation and cloth folding. Our results demonstrate that the synthesized skill library enables the system to transfer to novel tasks with decreasing human intervention, providing a steerable and data-efficient alternative to black-box robot learning. Website: https://robo-architect.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑