arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

引导动作流:面向流匹配视觉-语言-动作策略的Q引导推理

Guided Action Flow: Value-Guided Sampling for Frozen Vision-Language-Action Policies

Liuhaichen Yang, Zhuang Jiang, Chenchao Sheng, Ningwei Bai, Qichen Yin, Hanbo Ma, Junkai Liu, Junkai Sun, Dongcheng Lyu, Yi Dong, Zezhi Tang

arXiv 2607.02092首次发表:更新:

发表机构

Department of Computer Science, University College London; Department of Mechanical Engineering, University College London(计算机科学系,伦敦大学学院; 机械工程系,伦敦大学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出Guided Action Flow框架,通过冻结预训练策略并利用学习到的动作块评论家引导逆向流采样,在LIBERO任务中单任务评论家将成功率从68.0%提升至82.0%。

AI 中文摘要

流匹配视觉-语言-动作策略通过迭代传输过程生成机器人动作块,为无需重新训练基础策略的测试时引导创造了机会。我们在Guided Action Flow中研究这一机会,这是一个推理时框架,保持预训练的SmolVLA策略冻结,并使用学习到的动作块评论家来引导其逆向流采样器。该评论家从真实的成功和失败轨迹中训练,可以基于冻结的SmolVLA语言路径中的任务描述特征进行条件化,并且在采样过程中仅通过动作梯度使用。我们在LIBERO操作任务上评估该方法。单任务评论家将一个种子窗口的成功率从68.0%提高到82.0%,另一个从82.0%提高到86.0%。多家族任务描述评论家将验证成功率从46.0%提高到56.0%,而锁定的保留测试集增益为正但较小,从65.0%提高到67.5%。这些结果支持了冻结流匹配VLA策略的Q引导推理的可行性,同时表明评论家泛化和不确定性感知引导仍然是主要瓶颈。

英文摘要

Reinforcement learning can improve vision-language-action (VLA) policies beyond supervised fine-tuning, although this typically involves further updates to the policy parameters. For flow-matching policies, iterative action generation provides an additional opportunity to incorporate task information during inference. We introduce Guided Action Flow (GAF), which learns a compact, observation-conditioned action-value critic from robot task rollouts and applies its action gradient to steer reverse-time flow sampling. The supervised-fine-tuned VLA remains frozen throughout critic learning and deployment. Physical-robot experiments show an increase in aggregate success from 60.0% to 82.5% across six nominal manipulation tasks. Under six altered-lighting and object-distractor conditions evaluated on three of these tasks, aggregate success improves from 34.2% to 49.2%. Ablations and rollout analyses support the importance of the learned guidance direction and the critic's visual and proprioceptive inputs. With approximately 2.735M trainable critic parameters alongside a 0.45B-parameter VLA, GAF enables task outcomes to inform action generation through a compact inference-time guidance module.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑