Path-Coupled Bellman Flows for Distributional Reinforcement Learning
路径耦合贝尔曼流用于分布式强化学习
机构 * University of Science and Technology of China(中国科学技术大学)
AI总结 本文提出路径耦合贝尔曼流(PCBF),一种连续时间的分布式强化学习方法,通过学习回报分布的流匹配来解决现有方法在边界不匹配和高方差-bootstrap问题,实验表明其在分布保真度和训练稳定性方面有所提升。
Comments Accepted to the 43rd International Conference on Machine Learning (ICML 2026)
Journal ref Proceedings of the 43rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026