悲观极小极大学习用于单侧覆盖下的公私信息博弈
Pessimistic Minimax Learning for Public-Private Information Games under Unilateral Coverage
- Massachusetts Institute of Technology(麻省理工学院)
- Purdue University(普渡大学)
- California Institute of Technology(加州理工学院)
- University of Virginia(弗吉尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出悲观极小极大学习框架,用于公私信息博弈的离线均衡学习,引入单侧规定可集中性,并开发悲观算法及PPA-PMD框架,实现$\tilde{O}(1/\sqrt{n})$和$\tilde{O}(1/\sqrt{n}+1/\sqrt{T})$的利用率,为相关领域提供首个理论框架。
AI中文摘要:
我们研究具有公共和私人信息的两人零和情境博弈中的离线学习,其动机来自诸如拍卖和具有私人估值的谈判等战略环境。我们引入了单侧规定可集中性,并表明不对称信息可以通过其对均衡行为的影响来改变离线覆盖。对于有限状态-动作空间,我们开发了一种悲观算法,其利用率为$\tilde{O}(1/\sqrt{n})$,与完全观测的极小极大博弈的标准样本规模依赖性相匹配。我们进一步开发了一种悲观策略镜像下降框架PPA-PMD,用于一般函数逼近,并获得了统一的$\tilde{O}(1/\sqrt{n} + 1/\sqrt{T})$利用率,且参与者更新具有无遗憾特性。总之,这些结果为在公私信息约束下的离线均衡学习提供了第一个理论框架。
英文摘要:
We study offline learning in two-player zero-sum contextual games with public and private information, motivated by strategic settings such as auctions and negotiations with private valuations. We introduce unilateral prescriptive concentrability and show that asymmetric information can change offline coverage through its effect on equilibrium behavior. For finite state-action spaces, we develop a pessimistic algorithm with an $\tilde{O}(1/\sqrt{n})$ exploitability rate, matching the standard sample-size dependence for fully observed minimax games. We further develop a pessimistic policy mirror descent framework, PPA-PMD, for general function approximation and obtain a unified $\tilde{O}(1/\sqrt{n} + 1/\sqrt{T})$ exploitability rate with no-regret actor updates. Together, these results provide the first theoretical framework for offline equilibrium learning under public-private information constraints.