发表机构
Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于细化的流策略优化(RFPO),通过Q引导样本细化与自目标流匹配交替训练,解决流策略在线强化学习中隐式分布采样难题,在连续控制与多模态分布任务上优于基线。
AI 中文摘要
基于流的策略为在线强化学习提供了一种富有表现力的表示,但传统的流匹配要求从待建模的分布中采样。当期望的动作分布仅由Q函数隐式定义时,这构成了挑战,因为直接从所得分布中采样动作通常是难以处理的。我们提出了基于细化的流策略优化(RFPO),一种通过在Q引导的样本细化和自目标流匹配之间交替来训练在线强化学习中的流策略的新框架。RFPO首先使用当前流策略从高斯噪声生成动作,然后使用有限步随机细化过程将它们移向由Q函数诱导的基于能量的分布。每个细化后的动作随后与其对应的初始噪声样本配对,并用作流匹配训练的固定目标。通过反复细化自身输出并从所得目标中学习,RFPO将Q引导纳入策略,而无需直接从目标分布中采样,同时保留表示多种动作模式的能力。我们进一步对RFPO诱导的分布动态进行了理论分析。在六个连续控制任务中,RFPO在几乎所有任务上匹配或优于标准高斯策略基线。在具有不同几何形状的六个合成二维目标分布上的实验表明,RFPO能够捕获复杂的多模态结构而不会出现模式崩溃。
英文摘要
Flow-based policies offer an expressive representation for online reinforcement learning, but conventional flow matching requires samples drawn from the distribution to be modeled. This poses a challenge when the desired action distribution is defined only implicitly by a Q-function, since directly sampling actions from the resulting distribution is generally intractable. We propose Refinement-Based Flow Policy Optimization (RFPO), a novel framework for training a flow policy in online reinforcement learning by alternating between Q-guided sample refinement and self-target flow matching. RFPO first generates actions from Gaussian noise using the current flow policy and then uses a finite-step stochastic refinement procedure to move them toward an energy-based distribution induced by the Q-function. Each refined action is then paired with its corresponding initial noise sample and used as a fixed target for flow-matching training. By repeatedly refining its own outputs and learning from the resulting targets, RFPO incorporates Q-guidance into the policy without requiring direct samples from the target distribution, while retaining the capacity to represent multiple action modes. We further provide a theoretical analysis of the distributional dynamics induced by RFPO. Across six continuous-control tasks, RFPO matches or outperforms a standard Gaussian-policy baseline on almost every task. Experiments on six synthetic two-dimensional target distributions with diverse geometries demonstrate that RFPO captures complex multimodal structure without mode collapse.