arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

具有探索性策略的N人随机微分博弈的连续时间强化学习

Continuous-Time Reinforcement Learning for $N$-Player Stochastic Differential Games with Exploratory Policies

Jisheng Liu

arXiv 2607.19928首次发表:更新:

发表机构

School of Mathematical Sciences, Fudan University(复旦大学数学科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究N人非合作随机微分博弈的连续时间强化学习,参与者采用熵正则化探索策略,证明自然均衡概念与兼容性等价并建立条件,给出解耦对称博弈均衡情况,还介绍不成立时的构造及框架扩展与渐近速率。

AI 中文摘要

我们研究N人非合作随机微分博弈的连续时间强化学习。每个参与者采用熵正则化探索策略;给定其他参与者的行动,最优响应是吉布斯分布,纳什均衡要求这N个条件分布联合兼容。我们证明自然均衡概念——同时哈密顿最大化——等同于这种兼容性,并建立了一个充要条件,以最优q函数上可计算的交叉偏导数准则表示。对于解耦和对称博弈,纳什均衡无条件存在。当兼容性不成立时,坐标路径积分构造产生一个近似相关均衡,具有明确的二次KL散度界,当探索权重γ→∞时局部均匀消失。N人博弈的q函数框架扩展了[21]中的单智能体q学习理论,弱鞅特征推动了无模型的在线和离线算法。该框架扩展到遍历(无限期)设置,具有相同的局部均匀O_R(1/γ)渐近速率。

英文摘要

We study entropy-regularized exploratory control in finite $N$-player stochastic differential games under a response model in which each player conditions on the opponents' currently realized actions and evaluates continuation with that profile frozen. The resulting Gibbs best responses form a system of full conditional densities, which need not admit a common joint law. We characterize joint realizability by a cross-partial condition on the entropy-scaled Hamiltonian gradients and, on simply connected action domains, by an equivalent entropy-weighted potential structure. When compatibility fails, a coordinate-path construction yields a joint density whose full conditionals satisfy explicit quadratic Kullback--Leibler bounds. We extend the analysis to stationary discounted problems and derive martingale and policy-improvement characterizations for learning the frozen response maps. A two-player linear-quadratic example illustrates the compatibility criterion and the associated learning procedure.

Comments39 pages,4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑