发表机构
Indian Institute of Management Bangalore(印度管理学院班加罗尔分院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过Stackelberg博弈模型分析AI安全中反馈驱动的防御充分性,提出最小成本防御分配策略,并证明在持续发现与有效修复条件下可达到概率1的防御,且快速修复能缩短妥协时间。
AI 中文摘要
自动化测试、人工红队测试和事件响应的反馈可以在发现失败并有效修复时增强AI系统的防御能力。我们研究了这一反馈过程何时能提供充分保护,以及投资于该过程在经济上是否值得。我们首先证明,如果每个未解决的攻击都有持续被发现的机会,修复有效,且后续更新保留早期保护,那么由有限数量输入组成的攻击面以概率1得到防御。我们推导了完成时间界限,并将分析扩展到增长中的攻击面、跨相关攻击泛化的修复以及多种发现机制。这些结果区分了针对每个固定攻击的最终保护与在单一时间点的完全保护。然后,我们构建了一个防御者主导的Stackelberg博弈,其中防御者投资于主动发现和被动修复,并预判攻击者的搜索努力选择。我们刻画了阻止攻击的最小成本分配,以及防御者资助两种能力、仅一种能力或两者都不资助的均衡状态。数值实验展示了这些状态,并表明更快的修复可以在不减少妥协可能性的情况下缩短妥协持续时间。本理论还涉及生成语言模型中的性能限制。
英文摘要
Feedback from automated testing, human red teaming, and incident response can strengthen an AI system's defenses when discovered failures lead to effective repairs. We study when this feedback process provides sufficient protection and when investing in it is economically worthwhile. We begin by showing that an attack surface composed of finite number of inputs is defended with probability 1 if every unresolved attack has a persistent chance of discovery, repairs are effective, and subsequent updates preserve earlier protection. We derive completion-time bounds and extend the analysis to growing attack surfaces, repairs that generalize across related attacks, and multiple discovery mechanisms. These results distinguish eventual protection against each fixed attack from complete protection at a single time. We then formulate a defender-led Stackelberg game in which the defender invests in proactive discovery and reactive repair, anticipating the attacker's choice of search effort. We characterize the least-cost allocation that deters attack and the equilibrium regimes in which the defender funds neither capability, one capability, or both. Numerical experiments illustrate these regimes and show how faster repair can reduce compromise duration without reducing compromise probability.
Comments27 pages, 3 figures, 1 table