arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13755math.OC

具有部分不对称信息的零和动态博弈的最佳响应动力学

Best Response Dynamics for Zero-Sum Dynamic Games with Partial-Asymmetric Information

Yuxiang Guan, Iman Shames, Tyler H. Summers

AI总结:

针对部分不对称信息下的零和随机线性二次动态博弈,本文提出基于最佳响应动力学的方法,推导纯线性动态输出反馈控制策略的最佳响应显式表达式,实验表明博弈值迭代几次后收敛,还分析了相关因素对线性二次追逃博弈值的影响。

AI中文摘要:

本研究探讨一类在部分及不对称信息下的零和随机线性二次动态博弈(LQDGs)。信息不对称带来了与信念表示和心智理论相关的根本挑战,博弈参与者必须推断其他参与者的信念状态与估计值以制定策略。现有研究表明,将类似动态规划的分解方法应用于此类问题存在难度。本文提出一种基于最佳响应动力学的替代方法,为应对信念表示和心智理论挑战提供了思路;推导了纯线性动态输出反馈控制策略类别中各参与者最佳响应的显式表达式,其中每个控制器的内部状态维度为系统状态维度的整数倍。随着参与者迭代更新其最佳响应,会形成阶数越来越高的信念状态,进而产生无穷维内部状态。但数值结果显示,该博弈的值仅在几次迭代后便收敛,表明高阶信念状态带来的收益可忽略不计。本文还开展数值实验,分析了不对称信念、信念阶数、相对可控性与可观性以及直接前馈对线性二次追逃博弈值的影响。

英文摘要:

This work studies a class of zero-sum stochastic linear quadratic dynamic games (LQDGs) under partial and asymmetric information. Information asymmetry introduces fundamental challenges related to \textit{belief representation} and \textit{theory of mind}, where players must impute belief states and estimates of other players to inform their strategies. Existing work highlights the difficulty of applying dynamic programming-like decomposition approach to these problems. An alternative approach based on \textit{best response dynamics} is proposed, which provides insights into belief representation and theory of mind challenges. Explicit expressions for each player's best response within the class of pure linear dynamic output feedback control strategies are derived, where the internal state dimension of each control is an integer multiple of the system state dimension. As players iteratively update their best responses, they form increasingly higher-order belief states, leading to infinite-dimensional internal states. However, numerical results reveal that the game's value converges after only a few iterations, suggesting that higher-order belief states provide vanishing benefit. This work further conducts numerical experiments to analyze the impact of asymmetric beliefs, belief orders, relative controllability and observability, and direct feed-through on a linear quadratic pursuit-evasion game's value.

↑