arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GR(1)博弈的自适应策略

Adaptive Strategies for GR(1) Games

Shankaranarayanan Krishna, Kaushik Mallik, Abhilasha Sharma Suman

arXiv 2608.28391首次发表:更新:

发表机构

IIT Bombay; IMDEA Software Institute(印度理工学院孟买分校; IMDEA软件研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对GR(1)博弈提出自适应框架,将环境玩家视为非对抗智能体,通过监控活性属性调整策略,解决传统静态策略的保守性问题,其随机策略可渐近收敛至最优确定性策略,原型性能优于现有技术。

AI 中文摘要

我们考虑图上的两人GR(1)博弈,其中系统玩家Eve必须在环境玩家Adam面前满足\\[ \Box\Diamond A_1\land\cdots\land\Box\Diamond A_m \\;\implies\\; \Box\Diamond G_1\land\cdots\land\Box\Diamond G_n \\]。此处$A_1,\ldots,A_m$是对环境的假设,$G_1,\ldots,G_n$是系统必须提供的保证,而$\Box\Diamond S$表示“最终永远满足S”。传统静态策略过于保守:它们可能主动违反假设以简单满足蕴含关系,或在任何假设被违反时放弃所有保证。现有防止此类行为的方法会产生双指数级的膨胀。我们引入一种自适应框架,将Adam视为具有未知目标的非对抗性智能体。Eve监控Adam实际满足的假设,并在运行时调整策略以最大化满足的保证。我们方法的核心是一种用于监控活性属性$\Box\Diamond S$的新算法,使Eve能够实时估计哪些假设将被满足。Eve预先计算针对不同假设子集的最优策略,基于监控器输出动态调整对这些策略的概率分布。我们证明,当假设被违反时,Eve的随机自适应策略会渐近收敛到最大化保证的确定性策略。原型系统展示了其有效性和与现有技术相比更优的计算性能。

英文摘要

We consider two-player GR(1) games on graphs, where the system player Eve must satisfy \[ \Box\Diamond A_1\land\cdots\land\Box\Diamond A_m \;\implies\; \Box\Diamond G_1\land\cdots\land\Box\Diamond G_n \] against the environment player Adam. Here $A_1,\ldots,A_m$ are assumptions on the environment, $G_1,\ldots,G_n$ are guarantees the system must provide, and $\Box\Diamond S$ denotes ``always eventually $S$''. Traditional static strategies are overly conservative: they may actively violate assumptions to trivially satisfy the implication, or abandon all guarantees when any assumption is violated. Existing methods to prevent such behaviors incur doubly exponential blowup. We introduce an adaptive framework treating Adam as a non-adversarial agent with unknown objectives. Eve monitors which assumptions Adam actually meets and adapts her strategy at runtime to maximize satisfied guarantees. Central to our approach is a novel algorithm for monitoring liveness properties $\Box\Diamond S$, enabling Eve to maintain real-time likelihood estimates of which assumptions will be fulfilled. Eve pre-computes strategies optimal for different assumption subsets, deploying a probability distribution over them that dynamically adjusts based on monitor outputs. We prove that when assumptions are violated, Eve's randomized adaptive strategy converges asymptotically to the deterministic strategy maximizing guarantees. A prototype demonstrates effectiveness and superior computational performance compared to the state of the art.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑