arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

最优传输遇上强化学习:综述

Optimal Transport Meets Reinforcement Learning: A Survey

Yujie Zhu, Charles A. Hepburn, Matthew Thorpe, Giovanni Montana

arXiv 2610.01413首次发表:更新:

发表机构

University of Warwick(华威大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本综述系统梳理最优传输在强化学习目标与算法中的应用,分类现有方法,探讨动机、实际考量与开放问题,为相关研究提供全面参考。

AI 中文摘要

强化学习(RL)算法经常比较概率分布,例如由策略和专家诱导的状态访问分布、来自学习策略和离线数据集的动作分布,或来自学习模型和环境的转移分布。然而,当这些分布重叠较弱时,常用的散度可能变得无效,这种情况在模仿学习、离线RL和分布偏移下的部署中经常遇到。最优传输(OT)通过在地成本(该成本编码任务几何)下测量将概率质量从一个分布“移动”到另一个分布的成本,提供了一种替代方案。本综述涵盖了OT如何在RL目标和算法中使用。对于每种方法,我们识别:OT所扮演的角色、所比较的分布、所使用的OT公式以及时间结构的处理。除了对现有方法进行分类外,我们讨论了不同OT选择背后的动机、实际考虑因素(如成本设计和计算挑战),并强调了开放问题,包括可扩展的轨迹级传输、质量不匹配的原则性处理以及OT正则化RL的理论分析。

英文摘要

Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and experts, action distributions from learned policies and offline datasets, or transition distributions from learned models and environments. However, commonly used divergences may become ineffective when these distributions overlap weakly, which is frequently encountered in imitation learning, offline RL, and deployment under distribution shift. Optimal transport (OT) offers an alternative by measuring the cost of \emph{moving} probability mass from one distribution to another under a ground cost that encodes task geometry. This survey covers how OT is used inside RL objectives and algorithms. For each method, we identify: the role OT plays, the distributions compared, the OT formulation used, and the treatment of temporal structure. Beyond categorising existing methods, we discuss the motivations behind different OT choices, practical considerations such as cost design and computational challenges, and highlight open problems including scalable trajectory-level transport, principled handling of mass mismatch, and theoretical analysis for OT-regularised RL.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑