arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有确定性策略的扩展平均场控制的演员-评论家学习

Actor-Critic Learning for Extended Mean Field Control with Deterministic Policies

Ziheng Cheng, Xin Guo, Huyên Pham, Yufei Zhang

arXiv 2607.11005首次发表:更新:

发表机构

University of California, Berkeley; Ecole Polytechnique, CMAP; Imperial College London(加州大学伯克利分校; 巴黎政治经济学院(École Polytechnique); 伦敦帝国学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对连续时间扩展平均场控制问题,提出无模型强化学习框架,采用确定性反馈策略,建立灵敏度公式并推导策略梯度公式,经细化后得到含相关导数项的策略梯度,结合多种方法形成算法,数值实验验证了该方法的有效性。

AI 中文摘要

本文为连续时间扩展平均场控制问题开发了一个无模型强化学习框架,其中动力学和奖励可能依赖于状态和控制的联合分布。我们采用确定性反馈策略,在此策略下,状态-动作分布直接作为状态律的前推诱导得出。这避免了对随机核的优化,并绕过了扩展平均场设置中现有方法的关键限制。我们首先为参数化的麦克凯恩-弗拉索夫动力学建立了一个无模型灵敏度公式,并使用它推导出一个通过瓦瑟斯坦空间上的优势率函数表示的确定性策略梯度公式。然后,我们通过引入依赖于状态、动作和联合状态-动作分布的局部值和优势率表示来细化这个公式,得到一个包含动作导数和关于控制分布的测度导数项的策略梯度。这些特征导致了基于鞅的学习原理,并激发了一种结合粒子近似、依赖测度的神经网络、时间差分学习以及在动作或参数空间中探索的连续时间深度确定性策略梯度算法。在随机库克尔-斯梅尔共识控制和具有交易拥挤的最优清算方面的数值实验证明了所提出方法的效率、稳定性和鲁棒性,包括明确依赖控制分布的问题。

英文摘要

This paper develops a model-free reinforcement learning framework for continuous--time extended mean field control problems, where both the dynamics and reward may depend on the joint distribution of states and controls. We adopt deterministic feedback policies, under which the state--action distribution is induced directly as a push--forward of the state law. This avoids optimization over stochastic kernels and bypasses key limitations of existing approaches in extended mean field settings. We first establish a model--free sensitivity formula for parameterized McKean--Vlasov dynamics and use it to derive a deterministic policy gradient formula expressed through an advantage--rate function on the Wasserstein space. We then refine this formula by introducing local value and advantage--rate representations that depend on the state, action, and joint state--action distribution, yielding a policy gradient that includes both action derivatives and measure--derivative terms with respect to the control distribution. These characterizations lead to a martingale--based learning principle and motivate a continuous--time deep deterministic policy gradient algorithm combining particle approximations, measure--dependent neural networks, temporal--difference learning, and exploration in either action or parameter space. Numerical experiments on stochastic Cucker--Smale consensus control and optimal liquidation with trade crowding demonstrate the efficiency, stability, and robustness of the proposed method, including problems with explicit dependence on the control distribution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑