arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19399cs.GTcs.AI

网络安全博弈中的高效纳什均衡计算

Efficient Nash Equilibrium Computation for Cybersecurity Games

Michael Lanier, David Farmer, Yevgeniy Vorobeychik

首次发表
浏览论文内容

中文总结 AI 辅助

针对仿真型网络安全博弈中PSRO的收益估计瓶颈,提出RWPS预算估计器,仅模拟敏感单元并利用替代模型,通过实例相关界与覆盖结果提升效率,实验显示在多个博弈中达到更低利用度。

中文摘要 AI 辅助

在基于策略空间响应预言机(PSRO)的仿真型网络安全博弈中计算纳什均衡,其瓶颈在于收益估计:每个收益矩阵条目都需要对慢速模拟器进行蒙特卡洛回放,而策略和受限博弈求解则成本低廉。我们引入了后悔加权收益采样(RWPS),这是一种预算受限的估计器,仅模拟均衡所敏感的单元,并用在运行过程中先前模拟的每个条目训练出的替代模型填充其余单元。上确界范数误差界无法评估此类估计器,因为它由故意保留不精确的单元决定。我们证明了一个依赖于实例的界,该界以对手的均衡混合策略对误差进行加权,这是一个仅凭模拟数据即可计算的证书,以及一个覆盖结果,表明一旦与偏差相关的集合被模拟,替代误差便无法影响任一参与者的后悔值。在三个21x21的一般和博弈(两个合成博弈和一个非对称的Colonel Blotto博弈)中,细化后的界在估计器自身输出上紧致了四到六倍,且覆盖结果能提前预测哪些博弈成本低:小支撑博弈为矩阵的18%,而Blotto博弈为82%。在增长池PSRO中,RWPS在匹配预算下达到的利用度低于最小后悔优先搜索、信息增益搜索和渐进采样,并且在CyGym和ANSG网络模拟器上,它在最小预算下表现最优。

英文摘要

Game-theoretic analyses of cyber defence often compute equilibria of games whose payoffs exist only as the output of a simulator. Iterative equilibrium-finding methods grow a set of attacker and defender policies and need the payoff of every attacker--defender pair, so they are bottlenecked by payoff estimation: each payoff costs many simulator runs. We introduce Regret-Weighted Payoff Sampling (RWPS), which spends a fixed simulation budget on the payoffs the equilibrium actually depends on and predicts the rest with a model trained on every payoff measured so far. Standard error bounds for estimated games are driven by the worst-estimated payoff, so they cannot credit an estimator that is inaccurate only where accuracy does not matter. We prove a bound that weights payoff errors by the opponent's equilibrium strategy, a certificate that can be computed from simulated payoffs alone, and a condition under which errors in the predicted payoffs cannot change either player's regret. On three synthetic general-sum games, one of them a Colonel Blotto game of military resource allocation, the new bounds are four to six times tighter than the standard one, and RWPS finds less exploitable equilibria than minimum-regret-first search, information-gain search and progressive sampling at the same budget. On two cyber-defence simulators, CyGym and a new game whose hosts are LLM agents exposed to prompt injection, it gives the least exploitable equilibria at the smallest budgets.

发表机构

  • Washington University in St. Louis(圣路易斯华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

↑