AI 中文总结
针对时变攻击强度的Stackelberg安全博弈,结合混合整数线性规划神谕与相关算法,分别在全信息和赌博机反馈下获得对应遗憾界,经仿真验证方法有效鲁棒。
AI 中文摘要
本研究探讨时变攻击强度下重复Stackelberg安全博弈中的无遗憾在线学习问题。我们构建了扩展安全博弈模型,其中攻击者可选择多个目标,并在乐观平局打破规则下推导得到精确混合整数线性规划神谕。在全信息反馈下,该神谕与Follow-the-Perturbed-Leader算法结合,针对具有时变追随者数量、攻击强度和攻击者类型的非预知序列,可获得期望O(√T)的遗憾值;在多追随者共享固定攻击者类型的赌博机反馈场景中,我们利用重心跨度构造从聚合攻击观测中重构效用估计,得到期望O(T^(2/3))的遗憾值。大量仿真验证了该方法在全信息与部分信息反馈下的鲁棒性与有效性。
英文摘要
This work investigates no-regret online learning for Repeated Stackelberg Security Games (RSSGs) against attackers with time-varying attack intensities. Standard RSSG models often assume fixed attack intensities, limiting their ability to capture evolving attack capabilities. To bridge this gap, we formalize a model in which follower numbers, types, and attack intensities can vary across rounds. We derive an exact mixed-integer linear programming oracle under optimistic tie-breaking. Under full-information feedback, combining this oracle with online learning algorithm yields expected $\mathcal{O}(\sqrt{T})$ regret against adversarial attacker sequences. Under bandit feedback, where only aggregate attack is observed, we represent the payoffs of all candidate defense strategies through a small set of basis strategies. This representation enables unbiased estimation of all candidate strategies' payoffs from a single aggregate payoff per round. Building on this representation, we combine barycentric spanner with EXP2 to achieve expected $\mathcal{O}(\sqrt{T})$ regret against adversarial sequences. This guarantee contains no additional logarithmic factors in $T$ and matches an $Ω(\sqrt{T})$ lower bound established within our model. Numerical simulations demonstrate the effectiveness of the proposed algorithms.