arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无完美记忆的部分观测平均场博弈:最优性条件与均衡

Partially Observed Mean Field Games Without Perfect Recall: Optimality Conditions and Equilibria

Xuanping Zhang, Xiao Zhang, Wang Yao

arXiv 2609.00880首次发表:更新:

发表机构

Beihang University; Hangzhou International Innovation Institute of Beihang University(北京航空航天大学; 北京航空航天大学杭州国际创新研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究无完美记忆的部分观测平均场博弈,通过参数化环境、推导随机最大值原理、证明均衡存在性,并以银行间拆借为例对比了完美记忆与无完美记忆下的均衡行为差异。

AI 中文摘要

本文研究无完美记忆(WPR)的部分观测平均场博弈。代表性智能体观测到带噪信号,但时刻\textit{t}的控制仅使用\textit{G}^I_t=σ(y_t)——这是一个通常非嵌套的信息族。而条件总体分布则使用观测滤子\textit{F}^Y。这两个耦合层面依赖不同的信息尺度,难以在同一构造中闭合。我们通过状态、驱动变量和随机平均场项的确定性相容联合分布对环境进行参数化,从而在不扩大智能体控制信息的前提下保留其依赖结构。对于固定分布,参考测度与Girsanov定理推导出WPR随机最大值原理;所选响应由条件哈密顿量与WPR信念测度表示。隐状态与平均场项的联合路径后验给出条件总体分布的弱Kushner-Stratonovich表示。递归响应映射在相容分布的紧凸集上连续,因此由Schauder-Tychonoff定理可得弱WPR均衡。对于固定的均衡分布与反馈,相容的Yamada-Watanabe定理将路径唯一性提升为强实现。最后,通过一个线性二次银行间拆借示例对比完美记忆(PR)与WPR。PR响应遵循Kalman-Bucy反馈,而WPR响应需求解Fredholm-Volterra方程,且在高斯假设下关于当前观测是仿射的。数值实验展示了PR与WPR之间均衡行为的差异。

英文摘要

This paper studies partially observed mean field games without perfect recall (WPR). The representative agent observes a noisy signal, but the control at time \(t\) uses only \(\mathcal G_t^I=σ(y_t)\), a generally non-nested information family. The conditional population law instead uses the observation filtration \(\mathbb F^Y\). These coupled levels rely on different information scales and are difficult to close within one construction. We parameterize the environment by a deterministic compatible joint law of state, driving variables, and random mean field term, thereby preserving its dependence structure without enlarging the agent's control information. For a fixed law, a reference measure and Girsanov's theorem yield a WPR stochastic maximum principle; the selected response is represented by the conditional Hamiltonian and WPR belief measure. The joint path posterior of hidden state and mean field term gives a weak Kushner-Stratonovich representation of the conditional population law. A recursive response map is continuous on a compact convex set of compatible laws, so Schauder-Tychonoff yields a weak WPR equilibrium. For a fixed equilibrium law and feedback, a compatible Yamada-Watanabe theorem lifts pathwise uniqueness to a strong realization. Finally, a linear-quadratic interbank lending example compares perfect recall (PR) with WPR. The PR response follows the Kalman-Bucy feedback, whereas the WPR response solves a Fredholm-Volterra equation and is affine in the current observation under Gaussianity. The numerical experiment illustrates how equilibrium behavior differs between PR and WPR.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑