arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16889cs.ROcs.AIcs.CV

不要放弃BATON:通过智能体子任务探索和感知转换的记忆实现长程机器人操作

Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory

Bingxin Xu, Yuzhang Shang, Emilio Ferrara

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对长程机器人操作的误差累积与子任务转换问题,提出BATON方法,通过子任务探索与感知转换记忆,在RoboMemArena基准上提升任务成功率11.6%、累计成功率14.9%。

中文摘要 AI 辅助

长程机器人操作将多个接触密集型技能串联为多阶段任务。视觉-语言-动作(VLA)模型日益掌握单个技能,但串联后仍会失败:误差累积超出策略修正能力,且一个子任务会隐性约束后续子任务。一种有前景的方案是冻结VLA,交由大语言模型(LLM)智能体负责:它用语言规划、通过解析原语在自由空间移动、仅在接触密集型片段调用VLA、并将适配信息写入语言记忆。将其应用于长程任务时会出现两个问题:(1)能力来自测试时的全任务探索,其成本随阶段数呈乘法增长:若一个阶段需要T个回合,K阶段任务约需T^K个回合,且失败无法定位到具体阶段;(2)它无转换表示:VLA原语仅带有退出条件,无进入条件,因此子任务可能以后续子任务无法使用的形式完成。我们提出BATON。针对问题(1),BATON将子任务作为探索单元:每个子任务在成本较低的短程机制下探索,其解决方案存储于记忆中;长程轨迹由这些解决方案组合而成,而非整体发现,成本变为加法(T*K),且每个失败可归因于单个阶段。针对问题(2),BATON为探索配备感知转换的记忆:子任务内,验证智能体控制调用转换,仅在腕部视角确认场景就绪后才调用VLA;子任务间,交接转换会恢复被前序任务残留干扰的进入状态,前瞻转换会选择后续子任务可继承其结果的策略。全程不更新参数。在长程基准RoboMemArena上,BATON较现有最优方法(SoTA)将任务成功率提升11.6%,累计成功率提升14.9%。

英文摘要

Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VLA) and world-action models (WAMs) increasingly master individual skills, yet the chain still fails: errors compound beyond the policy's ability to correct, and one subtask silently constrains the next. A promising pathway freezes the VLA and puts an LLM coding agent in charge: it plans in language, moves in free space with analytic primitives, invokes the VLA only for contact-rich segments, and writes adaptation into language memory. Yet applied to long horizons, this recipe breaks twice. (1) Its competence comes from whole-task exploration at test time, whose cost is exponential in the number of stages: if one stage needs T episodes, a K-stage task needs on the order of T^K, and a failure does not reveal which stage caused it. (2) It has no representation of transitions: the VLA primitive carries an exit but no entry condition, and a subtask can succeed in a form its successor cannot use. We present BATON to address both failures. Against (1), BATON makes the subtask the unit of exploration: each subtask is explored in the cheap short-horizon regime and its solution stored in memory; a long-horizon trajectory is then composed from these solutions rather than discovered whole. Exploration cost becomes linear (KT), and each failure is attributed to one stage. Against (2), BATON equips exploration with a transition-aware memory. Within a subtask, a verifier agent governs the invocation transition: the VLA is invoked only after the wrist view confirms the scene is ready. Across subtasks, a handoff transition restores an entry state disturbed by the predecessor's residue, and a lookahead transition selects the strategy whose outcome the successor can inherit. On the RoboMemArena benchmark, BATON improves task success by 37.7% and cumulative success by 29.7% over the SoTA.

发表机构

  • University of Southern California(南加州大学)
  • University of Central Florida(中佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

↑