arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CAP-DO:用于跨相关零和博弈的认证双预言机求解的学习上下文动作提议

CAP-DO: Learned Contextual Action Proposals for Certified Double-Oracle Solving Across Related Zero-Sum Games

Mu Wang, Zhenkun Liu, Liang Liang, Guofu Zhang

arXiv 2607.24610首次发表:更新:

AI 中文总结

研究针对解决一系列相关零和博弈问题,提出CAP-DO框架,通过离线训练排序器、在线提议动作集,结合学习与当前博弈确定输出,理论上保留认证保证且有限收敛,实验中在多规模基准测试里有更好表现。

AI 中文摘要

许多安全和检查规划问题需要解决一系列相关的零和博弈。在这个序列中,可行的防御者和攻击者动作空间保持固定,而每个上下文通过目标值、检查有效性、成本和交互效应的变化诱导出不同的收益矩阵。双预言机(DO)通过迭代扩展受限博弈来解决大型零和博弈,而无需具体化完整的收益矩阵。然而,将标准DO独立应用于每个新的收益上下文需要从通用受限博弈重新开始搜索,并通过全空间最佳响应重新发现与上下文相关的动作。我们提出了上下文动作提议双预言机(CAP-DO),这是一个学习增强框架,用于为重复的上下文博弈热启动DO。离线时,CAP-DO从已解决的上下文中一次性训练单独的防御者和攻击者排序器。在线时,固定的排序器为每个新上下文提议初始受限动作集。学习因此确定认证搜索的起始位置,而当前博弈通过全空间最佳响应检查和双边证书仍可确定输出是否被接受。理论上,CAP-DO保留了DO的全博弈认证保证,所以每个被接受的输出都符合规定的证书容差。在标准精确预言机假设下,当扩展无上限时,CAP-DO也保持有限收敛。实验上,在一个非加法上下文检查博弈基准的三个规模上,每个玩家最多有9880个动作,在固定扩展预算下,CAP-DO平衡实现了更高的认证率,并且比冷启动、跟踪重用和启发式热启动使用更少的全空间最佳响应调用。

英文摘要

Many security and inspection-planning problems require solving a sequence of related zero-sum games. Across this sequence, the feasible defender and attacker action spaces re-main fixed, whereas each context induces a different payoff matrix through changes in target values, inspection effective-ness, costs, and interaction effects. Double Oracle (DO) solves large zero-sum games without materializing the full payoff matrix by iteratively expanding a restricted game. However, applying standard DO independently to each new payoff context requires restarting the search from a generic restricted game and requires rediscovering context-relevant actions through full-space best responses. We propose Con-textual Action Proposal Double Oracle (CAP-DO), a learning-augmented framework that warm-starts DO for repeated contextual games. Offline, CAP-DO trains separate defender and attacker rankers once from solved contexts. Online, the fixed rankers propose initial restricted action sets for each new context. Learning therefore determines where certified search starts, while the current game, through full-space best-response checks and a two-sided certificate, still determines whether the output is accepted. Theoretically, CAP-DO pre-serves DO's full-game certification guarantee, so every accepted output meets the prescribed certificate tolerance. Under standard exact-oracle assumptions, CAP-DO also retains finite convergence when expansion is uncapped. Empirically, across three scales of a non-additive contextual inspection-game benchmark, with up to 9,880 actions per player, CAP-DO-balanced achieves higher certification rates and uses few-er full-space best-response calls than cold-start, trace-reuse, and heuristic warm starts under fixed expansion budgets.

Comments10 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑