arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FBFM:一种用于世界-动作模型执行中流匹配的无训练异步反馈机制

FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution

Peize Li, Ruimeng Zhang, Ru Zhang, Cong Huang, Kai Chen, Shanghang Zhang

arXiv 2607.29235首次发表:更新:

发表机构

Peking University; Tsinghua University; Zhongguancun Academy; Zhongguancun Institute of Artificial Intelligence(北京大学; 清华大学; 中关村学院; 中关村人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对世界-动作模型分块反馈时间粒度粗、无法逐时间步纠错的问题,提出无训练异步反馈机制FBFM,在两类WAM上使LIBERO等任务成功率提升超5%,实现开环流生成与闭环真实世界动力学的桥接。

AI 中文摘要

尽管世界-动作模型(WAMs)通过在执行前预测视觉演化提升了机器人的长程控制能力,但长程可靠性要求对真实观测进行重复重对齐,而非递归滚动。现有WAMs通过在各块之间用真实数据刷新历史或键值(KV)缓存来解决这一问题,但这种分块反馈的时间粒度较粗,无法在单个时间步层面修正预测误差。为解决该问题,我们提出反馈流匹配(FBFM),这是一种无训练的推理机制,将重对齐操作嵌入主动生成的块内部。在流匹配过程中,FBFM对条件速度场应用掩码伪逆修正:它利用前一个动作块引导下一个动作块的生成,并使用执行该前一个块后观测到的图像引导下一帧预测。这种跨块配对(即来自一个块的反馈及时到达以影响下一个块)形成了异步循环,无需等待块边界即可修正误差。由于无训练,该机制提升了对意外事件的响应能力并抑制了长程任务中的漂移。我们在联合生成WAM(DreamZero)和分阶段WAM(LingBot-VA)上评估FBFM,在选定的LIBERO和RoboTwin2.0任务中,它在有利设置下将成功率提升了5%以上,真实世界机器人观测-预测诊断显示跟踪效果显著更好。我们认为FBFM为细粒度在线修正提供了新范式,将开环流生成与闭环真实世界动力学桥接起来。

英文摘要

Although world-action models (WAMs) enhance long-horizon robot control by predicting visual evolution before acting, long-horizon reliability demands repeated re-grounding in real observations--not recursive rollout. Existing WAMs address this by refreshing history or KV cache with ground-truth data between chunks. However, such chunk-wise feedback operates at a coarse temporal granularity and thus fails to correct prediction errors at the individual time-step level. To address this, we propose Feedback Flow Matching (FBFM), a training-free inference mechanism that pushes re-grounding inside the actively generated chunk. During flow matching, FBFM applies a masked pseudoinverse correction to the conditional velocity field: it leverages the preceding action chunk to guide generation of the next action chunk, and uses the image observed after executing that preceding chunk to guide the next frame prediction. This cross-chunk pairing--where feedback from one chunk arrives in time to shape the next--creates an asynchronous loop that corrects errors without waiting for chunk boundaries. Being training-free, the mechanism improves responsiveness to unexpected events and suppresses drift in long-horizon tasks. We evaluate FBFM on both a joint-generation WAM (DreamZero) and a stage-wise WAM (LingBot-VA). On selected LIBERO and RoboTwin2.0 tasks, it improves success rates by over 5% in favorable settings, and real-world robot observation-prediction diagnostics show notably better tracking. We argue that FBFM offers a new paradigm for fine-grained online correction, bridging open-loop flow generation with closed-loop real-world dynamics.

Comments29 pages, 5 figures. Preprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑