发表机构
Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VGFM提出一种离线强化学习框架,利用条件流匹配模型在动作空间中实现密集值引导,避免时间反向传播,在OGBench任务上展现高效且可扩展的策略学习性能。
AI 中文摘要
最近的机器人学习范式越来越依赖于大规模的离线机器人交互数据集来训练控制策略。富有表现力的生成模型能够实现丰富且多模态的动作表示,从而扩展了这一范式在复杂机器人控制中的能力。然而,使用多步生成式智能体进行策略改进仍然具有挑战性。在离线强化学习(RL)中,沿着生成轨迹引入基于值的目标通常会带来大量的训练复杂性,包括时间反向传播(BPTT)、辅助架构或蒸馏损失。我们提出了值引导的流匹配(VGFM),这是一个可扩展的离线RL框架,能够在基于流的策略中实现密集的值引导塑形,同时避免BPTT和额外的算法开销。VGFM将策略参数化为动作(x预测)空间中的条件流匹配模型,确保每个中间流步骤都能产生一个有效的机器人动作,该动作可以由标准的离线RL评论员直接评估。这种设计允许在随机采样的流时间点应用值引导,而无需对整个生成轨迹进行微分,同时通过改变底层流常微分方程(ODE)的离散化方式(无需重新训练)来保持推理时的灵活性。在OGBench中的机器人运动和控制任务上进行的评估表明,VGFM在严格的评估协议下,在广泛的任务中取得了强劲的性能。在极少超参数调整的情况下,这些结果表明VGFM为长时域、目标导向的机器人控制中的富有表现力的策略学习提供了一种简单、可扩展且有效的方法。
英文摘要
Recent robot learning paradigms increasingly rely on large offline datasets of robotic interactions to train control policies. Expressive generative models enable rich and multimodal action representations, expanding the capability of this paradigm for complex robotic control. However, policy improvement with multi-step generative actors remains challenging. In offline reinforcement learning (RL), incorporating value-based objectives along generative trajectories often introduces substantial training complexity, including backpropagation through time (BPTT), auxiliary architectures, or distillation losses. We propose Value-Guided Flow Matching (VGFM), a scalable offline RL framework that enables dense value-guided shaping within a flow-based policy while avoiding BPTT and additional algorithmic overhead. VGFM parameterizes the policy as a conditional flow-matching model in action (x-prediction) space, ensuring that each intermediate flow step produces a valid robot action that can be directly evaluated by a standard offline RL critic. This design allows value guidance to be applied at randomly sampled flow times without differentiating through the entire generative trajectory, while preserving inference-time flexibility by varying the discretization of the underlying flow ODE without retraining. Evaluated on robotic locomotion and manipulation tasks in OGBench, VGFM achieves strong performance across a wide range of tasks under rigorous evaluation protocols. With minimal hyperparameter tuning, these results demonstrate that VGFM provides a simple, scalable, and effective approach for expressive policy learning in long-horizon, goal-oriented robotic control.
CommentsIROS 2026