arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

未来锚定验证与在线恢复:面向世界动作模型

Future Anchored Verification and Online Recovery for World Action Models

Zhibin Qin, Zhenxiong Tan, Xinchao Wang

arXiv 2610.06280首次发表:更新:

发表机构

National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对世界动作模型执行偏离预测导致动作失效的问题,提出FAVOR框架,利用预测帧作为锚点进行验证与恢复,在不修改策略下提升LIBERO与LIBERO-Plus任务成功率。

AI 中文摘要

世界动作模型(WAMs)已成为机器人操作领域一种前景广阔的新范式。其工作方式为:首先预测任务应如何执行,然后从该预测的未来状态中解码出具体动作。然而,一旦实际执行偏离了预测,剩余的动作便会失效。仅从已处于分布外(out-of-distribution)的状态重新规划,往往无法恢复任务仍然需要的内容;现有的执行监控器仅能决定何时停止,却不能决定应恢复什么。我们观察到,答案其实已在手边:WAM在行动前预测出的未来帧,恰好描绘了它意图经过的状态序列。为此,我们提出FAVOR(未来锚定验证与在线恢复)——一个轻量级框架,将这些预测帧保留为锚点,并利用它们进行验证与恢复。锚点验证器(Anchor Verifier)将每个观测与对应锚点及已执行的动作进行比较,以标记出会破坏任务的偏差。锚点引导恢复(Anchor-Guided Recovery)则借助视觉-语言模型,将标记出的锚点转化为一条简短的纠正指令。在强化后的指令引导下,WAM执行该指令以返回至预期的未来状态,随后继续执行原任务。在不修改策略的前提下,FAVOR将基础WAM在LIBERO上的任务成功率从97.85%提升至98.10%,在LIBERO-Plus上从72.60%提升至72.98%。

英文摘要

World action models (WAMs) have emerged as a promising paradigm for robotic manipulation. They act by first predicting how a task should be performed and then decoding the actions from that future. However, the remaining actions are invalid once execution drifts from the prediction. Simply replanning from the already out of distribution state rarely restores what the task still requires; existing execution monitors decide when to stop, but not what to restore. We observe that the answer is already in hand: the future the WAM predicted before acting depicts exactly the states it intended to pass through. We introduce FAVOR (Future Anchored Verification and Online Recovery), a lightweight framework that keeps these predicted frames as anchors and uses them for verification and recovery. An Anchor Verifier compares each observation with its anchor, together with the executed actions, to flag deviations that break the task. Anchor-Guided Recovery uses a vision-language model to turn the flagged anchor into a short corrective instruction. Under strengthened instruction guidance, the WAM executes this instruction to return to the intended future. It then resumes the task. FAVOR raises the task success of the base WAM from 97.85% to 98.10% on LIBERO and from 72.60% to 72.98% on LIBERO-Plus without modifying the policy.

Comments18 pages, 5 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑