arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

历史依赖型SA-MDP中的ε-纳什均衡

Epsilon-Nash Equilibria in History-Dependent SA-MDPs

Brandon Gary Kaplowitz, Dominik Bohnet Zurcher, Akash Agrawal, Tala Jafari, Christian Schroeder de Witt, Paul W. Goldberg

arXiv 2609.18829首次发表:更新:

发表机构

Department of Engineering Science, University of Oxford; Department of Computer Science, University of Oxford(牛津大学工程科学系; 牛津大学计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究历史依赖型状态对抗马尔可夫决策过程,证明通用历史依赖均衡策略不存在,并提出计算初始状态依赖ε-纳什均衡的首个算法,通过约简为约束零和部分可观测随机博弈,在小型及Atari基准上验证了可扩展性。

AI 中文摘要

我们研究状态对抗马尔可夫决策过程(SA-MDP),将其视为观测空间攻击的博弈:在每一步中,智能体根据收到的观测选择一个动作,而知晓智能体真实状态的对手则在状态相关的邻近集合内选择一个扰动观测。现有工作侧重于马尔可夫策略,我们则为历史依赖下的SA-MDP开发了一种解概念和计算方法。这一动机源于结果表明历史依赖可以实质性地改变均衡结果,并迫使智能体和对手都调整其策略。首先,我们证明了不存在通用的(即对初始状态分布不可知的)历史依赖均衡策略。针对这一发现,我们的主要结果提出了计算初始状态依赖均衡的ε-近似值的首条算法路径。我们通过将SA-MDP约简为策略等价的约束零和单边部分可观测随机博弈来实现这一点。最后,我们在小型可解析验证的博弈上测试了我们的算法,并表明它能扩展到更大、更现实的基准,包括具有12期前瞻视野的Atari Freeway rollout。

英文摘要

We study state-adversarial Markov decision processes (SA-MDPs) as games of observation-space attacks: at each step, an agent selects an action from a received observation while an adversary$\unicode{x2014}$who knows the true state the agent is in$\unicode{x2014}$chooses a perturbed observation within a state-dependent proximity set. While existing work focuses on Markovian policies, we develop a solution concept and computational approach for SA-MDPs under history dependence. History dependence can materially change equilibrium outcomes and can force both the agent and the adversary to adapt their strategies. First, we prove the non-existence of universal (agnostic of the initial state distribution) history-dependent equilibrium policies. Our main result presents the first algorithmic route to computing $ε$-approximations of initial-state dependent equilibria. We do so by reducing SA-MDPs to a strategically equivalent constrained zero-sum one-sided partially observable stochastic game. We test our algorithm on small analytically verifiable games and show that it scales to larger, more realistic benchmarks, including Atari Freeway rollouts with a 12-period ahead horizon.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑