arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17940eess.SYcs.SY

含状态与控制依赖噪声的线性二次随机微分博弈的策略迭代

Policy Iteration for Linear-Quadratic Stochastic Differential Games with State- and Control-Dependent Noise

Karl Handwerker, Felix Thömmes, Lucas Günther, Balint Varga, Sören Hohmann

AI总结:

本文针对含状态与控制依赖噪声的线性二次随机微分博弈,提出序列PI算法与同伦初始化方法,推导弗雷歇导数闭式表达式,验证了算法及分析结果的有效性。

AI中文摘要:

本文提出一种针对含状态与控制依赖噪声的随机微分博弈的新型序列策略迭代(PI)方法,该方法的更新过程保持均方稳定性,从而保证迭代适定。我们进一步推导了序列PI映射在纳什均衡处的弗雷歇导数闭式表达式,所得刻画揭示了控制依赖噪声、策略评估灵敏度及更新顺序如何控制局部误差传播,并给出了局部线性收敛的显式充分条件。由于寻找初始稳定解是策略迭代中的主要挑战,我们还提出一种基于同伦的初始化方法,以确保获得有效起始点。通过数值例子验证了所提PI算法及分析结果的有效性。

英文摘要:

This paper presents a novel sequential policy iteration (PI) method for stochastic differential games with state- and control-dependent noise. The updates preserve mean-square stability, so that the iteration is well posed. We further derive a closed-form expression for the Fréchet derivative of the sequential PI map at a Nash equilibrium. The resulting characterization reveals how control-dependent noise, policy-evaluation sensitivity, and update ordering govern local error propagation, and yields explicit sufficient conditions for local linear convergence. Since finding an initial stabilizing solution is a major challenge in policy iteration, we also propose a homotopy-based initialization that ensures a valid starting point. The effectiveness of the proposed PI algorithm and the analytical results are verified through a numerical example.

↑