arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16978cs.ROcs.LG

VLCP:用于机器人操纵的视觉语言控制策略闭环代码重规划

VLCP: Vision Language Control Policy Closed-Loop Code Replanning for Robot Manipulation

  • University of Monastir(莫纳斯提尔大学)
  • St. Mildred’s-Lightbourn School(圣米尔德雷德-莱特本学校)
  • Cornell University(康奈尔大学)
  • Carnegie Mellon University(卡内基梅隆大学)
  • Silverstream AI(银流人工智能公司)
  • Mila -- Quebec AI Institute(米拉-魁北克人工智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

Dhia Naouali, Minghan Wu, Claudia Wong, Abhinav Puthran, Omar G. Younis

AI总结:

VLCP是一种无需训练的机器人操纵策略,通过在单回合内每K步由VLM重写控制代码闭合循环,在57项任务的评估中综合成功率达35.1%,较开环策略提升十倍。

AI中文摘要:

将前沿视觉语言模型(VLM)转化为机器人策略通常需要对其进行微调,使其输出预训练时从未见过的动作表示,这会丢失该模型值得选用的大部分推理能力。相反,我们保留VLM冻结,使其将策略编写为简短的Python控制函数,无需演示也无需微调。不过,一次性编写该代码属于开环操作。现有闭环方法存在反应层级不当的问题:它们会重试固定策略或选择不同子任务,但从不重写失败的代码。VLCP在单回合内于失败实际所在的控制代码层面闭合了循环。每K步,VLM会从多视角RGB、本体感受状态和状态增量中重新观测场景,再根据刚观测到的内容重写控制函数,从而在失败加剧前将其捕获。我们在57项任务的MuJoCo/RoboVerse套件上进行评估。这种无训练策略的综合成功率达35.1%,而每回合仅查询一次的相同系统仅为3.5%。该十倍差距在所有场景族中均存在不重叠的置信区间。该增益源于失败抓取的回合内恢复率达27.3%:开环控制器会将错失的抓取带到回合结束,而VLCP会在下次重规划时重新观测并修复。此外,该循环成本低廉:中位数84%的输入token命中缓存,单回合仅需约10次紧凑查询,且任意重规划期间编写的控制块会持续保存到跨回合技能库中,供后续提示词复用。

英文摘要:

Turning a frontier vision-language model into a robot policy usually means fine-tuning it to emit an action representation it never saw in pretraining, which throws away much of the reasoning that made the model worth reaching for. We go the other way and keep the VLM frozen. It writes the policy as a short Python control function, with no demonstrations and no fine-tuning. Writing that code once is open-loop, though. Existing closed-loop methods react at the wrong level: they retry a fixed policy or pick a different subtask, but never rewrite the code that failed. VLCP closes the loop where the failure actually lives, on the control code, within a single episode. Every $K$ steps the VLM re-observes the scene from multi-view RGB, proprioceptive state, and a state delta, then rewrites the control function from what it just saw, so a failure is caught before it compounds. We evaluate on a 57-task MuJoCo/RoboVerse sweep. This training-free policy reaches $35.1\%$ pooled success, against $3.5\%$ for the identical system queried once per episode. That tenfold gap holds with non-overlapping confidence intervals in every scene family. The gain traces to a $27.3\%$ within-episode recovery rate on failed grasps: a miss an open-loop controller would carry to the end of the episode gets re-observed and fixed at the next replan. And the loop stays cheap. A median $84\%$ of input tokens hit cache, an episode needs only about $10$ compact queries, and control blocks written during any replan persist to a cross-episode skill library reused in later prompts.

↑