arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10369cs.ROcs.AI

VINE:驯服强化学习中的生成控制策略

VINE: Taming Generative Control Policies for Reinforcement Learning

Rushuai Yang, Zhuo Han, Houlin Li, Hecheng Wang, Zhichao Wu, Rui Zhang, Zhaowei Zhang, Zihong Chen, Xiaohan Yan, Chiming Liu, Yi Chen, Wei Shan, Maoqing Yao

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对流匹配策略在值梯度强化学习中训练不稳定的问题,提出VINE方法。该方法通过在去噪步骤重建插值状态,创建稳定可微路径,实现稳定的端到端值梯度优化,在相关基准和任务中表现优于现有方法。

中文摘要 AI 辅助

流匹配策略已成为机器人学习中一种有效的策略参数化方法。它能从噪声中迭代生成动作,对复杂多模态动作分布进行高表达建模。然而,先前工作发现用值梯度强化学习扩展这些策略常导致训练不稳定。现有方法将其归因于迭代生成并避免端到端值梯度优化。本文表明不稳定并非源于迭代生成本身,而是源于最初为行为克隆设计的香草采样策略,在值梯度强化学习下变得脆弱。基于此,提出VINE,一种面向强化学习的采样方法,能为流匹配策略实现稳定的端到端值梯度优化。VINE在每次去噪步骤重建新的插值状态,创建稳定可微路径用于值梯度传播,同时与原始流匹配去噪过程兼容。结果,VINE在不牺牲端到端值梯度优化的情况下保留了流匹配的表达性和迭代生成能力。尽管通过所有十个去噪步骤进行端到端反向传播,VINE在OGBench离线强化学习基准和现实世界机器人操纵任务上实现了稳定的策略改进并持续优于现有强化学习方法。

英文摘要

Flow-matching policies have emerged as an effective policy parameterization for robot learning. They iteratively generate actions from noise, enabling highly expressive modeling of complex and multimodal action distributions. However, prior works observed that scaling these policies with value-gradient reinforcement learning (RL) often leads to training instability. Existing methods attribute this instability to iterative generation and therefore avoid end-to-end value-gradient optimization by sacrificing iterative generation, high expressiveness, or value-gradient optimization. Contrary to prior belief, we show the instability does not stem from iterative generation itself, but from the vanilla sampling strategy originally designed for behavior cloning, which becomes brittle under value-gradient RL. Motivated by this insight, we propose VINE, an RL-oriented sampling method that enables stable end-to-end value-gradient optimization for flow-matching policies. Instead of following a single flow trajectory, VINE reconstructs a new interpolation state at every denoising step, creating a stable differentiable path for value-gradient propagation while remaining compatible with the original flow-matching denoising process. As a result, VINE preserves the expressiveness and iterative generation of flow-matching without sacrificing end-to-end value-gradient optimization. Despite performing end-to-end backpropagation through all ten denoising steps, VINE achieves stable policy improvement and consistently outperforms state-of-the-art RL methods on the OGBench offline RL benchmark and real-world robotic manipulation task. Videos are available on our website: https://agibottech.github.io/vine.

发表机构

  • AgiBot(未知机构)
  • The Hong Kong University of Science and Technology(香港科技大学)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

↑