发表机构
University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出带转移不确定性的一般和并发随机博弈的鲁棒PAC学习框架,通过维护转移核置信集、引入纳什边际特征,在满足可达性条件下以多项式样本复杂度终止,基准博弈实验验证其性能与理论一致性。
AI 中文摘要
我们提出了首个针对带转移不确定性的一般和并发随机博弈(CSGs)的可能近似正确(PAC)学习框架,同时解决纳什均衡(NE)存在性的挑战。我们的算法维护基于数据的转移核的L¹置信集,并求解鲁棒CSG以计算社会福利最优的ε-纳什均衡(ε-NE),同时采用基于鲁棒马尔可夫决策过程(MDP)的探索机制来驱动联合状态-动作覆盖。关键是,我们引入了纳什边际特征,使得能够对均衡存在性进行有原则的推理:该框架要么返回社会福利值与最优值相差ε的ε近似NE,要么提供无精确NE存在的可靠证明。在相关状态-动作对上满足最小可达性条件p_reach>0时,算法在多项式数量的轨迹样本后终止,样本复杂度为Õ(R_max²H⁴|S|²|A|/(p_reachε²))。在基准CSGs上的实证结果表明,其性能接近最优,能正确处理均衡的存在与否,且样本复杂度与理论一致。
英文摘要
We introduce the first Probably Approximately Correct (PAC) learning framework for general-sum concurrent stochastic games (CSGs) with transition uncertainty, while addressing the challenge of Nash equilibrium (NE) existence. Our algorithm maintains data-driven $L^1$ confidence sets over transition kernels and solves a robust CSG to compute a social-welfare optimal $\varepsilon$-NE, using a robust MDP-based exploration mechanism to drive joint state-action coverage. Crucially, we introduce a Nash margin characterisation that enables principled reasoning about equilibrium existence: the framework either returns an $\varepsilon$-approximate NE whose social-welfare value is $\varepsilon$-close to optimal, or provides a sound certificate that no exact NE exists. Under a minimum reachability condition $p_{\mathrm{reach}}>0$ over relevant state-action pairs, the algorithm terminates after a polynomial number of trajectory samples, with sample complexity $\widetilde{O}\left( {R_{\max}^2 H^4 |S|^2 |A| / (p_{\mathrm{reach}} \varepsilon^2)} \right)$. Empirical results on benchmark CSGs demonstrate near-optimal performance, correct handling of equilibrium (non-)existence, and sample complexity consistent with theory.
CommentsCamera-ready version of a paper accepted to NeurIPS 2026. Main text: 10 pages, 1 figure, 2 tables; Appendix: 22 pages, 2 figures, 1 table. Minor revisions to the experiment compute resources