发表机构
ShanghaiTech University; Zhejiang University(上海科技大学; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究以防御者为中心评估越狱攻击,提出A - MESS框架,通过AttackSHAP估计攻击效用,在不同场景下实验发现攻击成功率排名与防御者效用弱对齐,有限查询能准确估计,直接优化子集安全效用更强,建议将越狱攻击视为安全改进资源。
AI 中文摘要
对大语言模型的越狱攻击通常以攻击者为中心的指标(如攻击成功率)来评估,但能攻破模型的攻击不一定对提升其安全性有用。我们提出以防御者为中心的越狱评估观点,即通过将攻击用作安全训练的红队数据时所带来的下游安全改进来评估。在此基础上,我们引入A - MESS(最小有效攻击子集选择),这是一个与设置无关的框架,用于从黑盒子集效用观察中归因和选择越狱攻击。A - MESS估计AttackSHAP,这是基于Shapley值的分数,可将边际效用归因于单个攻击,并通过贪婪或基于代理的优化在用户指定预算下选择紧凑攻击子集。在受控效用场景和真实大语言模型安全设置中,我们发现攻击成功率排名与以防御者为中心的效用弱对齐,有限效用查询能准确估计AttackSHAP,直接优化子集产生的安全效用比以攻击者为中心或仅归因选择更强。这些结果表明应将越狱攻击视为提升安全的资源,而非仅用于攻破模型的工具。
英文摘要
Jailbreak attacks on large language models are usually evaluated by attacker-centric metrics such as attack success rate (ASR), yet an attack that breaks a model is not necessarily useful for improving its safety. We propose a defender-centric view of jailbreak evaluation, where attacks are evaluated by the downstream safety improvements they enable when used as red-teaming data for safety training. Building on this view, we introduce A-MESS (Minimal Effective Attack-Subset Selection), a setting-agnostic framework for attributing and selecting jailbreak attacks from black-box subset utility observations. A-MESS estimates AttackSHAP, a Shapley-based score that attributes marginal utility to individual attacks and selects compact attack subsets under user-specified budgets via greedy or surrogate-based optimization. Across controlled utility landscapes and real LLM safety settings, we find that ASR rankings are weakly aligned with defender-centric utility, that AttackSHAP can be estimated accurately with limited utility queries, and that directly optimizing subsets yields stronger safety utility than attacker-centric or attribution-only selection. These results suggest evaluating jailbreak attacks as resources for improving safety, not only as tools for breaking models.