arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过值加权最优传输学习多模态一步流策略

Learning Multimodal One-step Flow Policy via Value-weighted Optimal Transport

Jaehun Shon, Jinha Choi, Jongwook Jeon, Jongmin Lee

arXiv 2609.15883首次发表:更新:

发表机构

Yonsei University(延世大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对离线强化学习中一步流策略学习的模式坍缩和分布外过估计问题,提出基于值加权最优传输的OptiFlow框架,通过传输引导实现分布内高值模态利用,在多个基准上表现优异。

AI 中文摘要

离线强化学习旨在仅从固定数据集中学习策略,这些数据集通常包含多模态动作分布。流策略能够自然地表示此类多模态行为,但学习高效的一步流策略仍然具有挑战性:标准的值引导常常导致模式坍缩或在分布外区域利用过估计偏差。为解决此问题,我们提出了通过最优传输的一步流策略(OptiFlow),一种将一步流策略学习视为结构化样本分配问题的框架。OptiFlow联合训练一个值感知的参考流策略和一个高效的一步策略,通过状态级熵正则最优传输耦合它们的动作样本。对于每个状态,评论家估计的值定义了蒸馏目标动作的优先级,而动作距离成本确保了几何上兼容的配对。通过避免直接最大化评论家,我们的传输引导方法通过将一步策略锚定到高值、数据集支持的模态上,实现了分布内利用,而无需担心分布外发散的风险。实验结果表明,OptiFlow有效地捕获了最优多模态行为,并在多种离线强化学习基准上取得了强劲的性能。我们的代码可在以下网址获取:此https URL。

英文摘要

Offline reinforcement learning aims to learn a policy solely from fixed datasets, which often contain multimodal action distributions. Flow policies can naturally represent such multimodal behaviors, but learning an efficient one-step flow policy remains challenging: standard value guidance often leads to mode collapse or exploits overestimation bias in out-of-distribution regions. To address this, we introduce One-step Flow policy via Optimal Transport (OptiFlow), a framework for one-step flow policy learning as a structured sample-allocation problem. OptiFlow jointly trains a value-aware reference flow policy and an efficient one-step policy, coupling their action samples through state-wise entropic optimal transport. For each state, critic-estimated values define the priority of distillation target actions, while the action-distance cost ensures geometrically compatible pairings. By avoiding direct critic maximization, our transport-guided approach enables in-distribution exploitation by anchoring the one-step policy to high-value, dataset-supported modes without the risk of out-of-distribution divergence. Experimental results demonstrate that OptiFlow effectively captures optimal multimodal behaviors and achieves strong performance across diverse offline RL benchmarks. Our code is available at https://github.com/Yonsei-DILLab/OptiFlow.

CommentsPreprint, 38 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑