发表机构
Department of Computer Science, University College London; Department of Mechanical Engineering, University College London(计算机科学系,伦敦大学学院; 机械工程系,伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Guided Action Flow框架,通过冻结预训练策略并利用学习到的动作块评论家引导逆向流采样,在LIBERO任务中单任务评论家将成功率从68.0%提升至82.0%。
AI 中文摘要
流匹配视觉-语言-动作策略通过迭代传输过程生成机器人动作块,为无需重新训练基础策略的测试时引导创造了机会。我们在Guided Action Flow中研究这一机会,这是一个推理时框架,保持预训练的SmolVLA策略冻结,并使用学习到的动作块评论家来引导其逆向流采样器。该评论家从真实的成功和失败轨迹中训练,可以基于冻结的SmolVLA语言路径中的任务描述特征进行条件化,并且在采样过程中仅通过动作梯度使用。我们在LIBERO操作任务上评估该方法。单任务评论家将一个种子窗口的成功率从68.0%提高到82.0%,另一个从82.0%提高到86.0%。多家族任务描述评论家将验证成功率从46.0%提高到56.0%,而锁定的保留测试集增益为正但较小,从65.0%提高到67.5%。这些结果支持了冻结流匹配VLA策略的Q引导推理的可行性,同时表明评论家泛化和不确定性感知引导仍然是主要瓶颈。
英文摘要
Reinforcement learning can improve vision-language-action (VLA) policies beyond supervised fine-tuning, although this typically involves further updates to the policy parameters. For flow-matching policies, iterative action generation provides an additional opportunity to incorporate task information during inference. We introduce Guided Action Flow (GAF), which learns a compact, observation-conditioned action-value critic from robot task rollouts and applies its action gradient to steer reverse-time flow sampling. The supervised-fine-tuned VLA remains frozen throughout critic learning and deployment. Physical-robot experiments show an increase in aggregate success from 60.0% to 82.5% across six nominal manipulation tasks. Under six altered-lighting and object-distractor conditions evaluated on three of these tasks, aggregate success improves from 34.2% to 49.2%. Ablations and rollout analyses support the importance of the learned guidance direction and the critic's visual and proprioceptive inputs. With approximately 2.735M trainable critic parameters alongside a 0.45B-parameter VLA, GAF enables task outcomes to inform action generation through a compact inference-time guidance module.