arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Fair ASR:在共享目标调用预算下重新评估黑盒越狱攻击

Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

Zhida He, Xiaoyu Wen, Han Qi, Ziyuan Zhou, Peng Yu, Jiajia Li, Chaochao Lu, Qiaosheng Zhang

arXiv 2608.17360首次发表:更新:

发表机构

Shanghai AI Laboratory; Fudan University; Shanghai Jiao Tong University(上海人工智能实验室; 复旦大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出Fair-ASR评估协议,重新评估11种黑盒越狱攻击,发现攻击排名随预算变化,进而提出ReCode攻击,在20次目标调用预算下于GPT-5实现85% ASR且效率优异。

AI 中文摘要

可靠的越狱评估对于评估大语言模型(LLM)的安全性至关重要,但现有大多数研究仅依赖攻击成功率(ASR),未考虑其对攻击预算的依赖性,导致不同方法间的比较不公平。现有的计算感知评估将异构资源简化为浮点运算次数(FLOPs),但这对于黑盒模型难以估算,且无法捕捉特定资源的约束。为提供可比较的评估基础,我们提出Fair-ASR,这是一种在共享目标调用预算B下针对黑盒越狱攻击的评估协议,使用目标调用作为直接可观测且与方法无关的比较轴,同时单独追踪攻击者调用以进行效率分析。我们在Fair-ASR协议下重新评估11种代表性攻击,发现攻击排名会随目标调用预算发生显著变化;在同等目标访问权限下,简单的随机扰动和手工模板仍具有很强的竞争力;且在目标调用和攻击者调用两方面,没有任何经评估的LLM驱动方法兼具高效性。受此效率差距的启发,我们提出ReCode,一种组合式高预算效率攻击,它结合脱敏重写与Fair-ASR识别出的两种高效低成本基元。在20次目标调用的预算下,ReCode在GPT-5上达到85%的ASR,同时平均每个请求仅需7.19次攻击者调用,在目标调用和攻击者调用两方面均表现出很强的效率。

英文摘要

Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) without accounting for its dependence on attack budgets, resulting in unfair comparisons across methods. Existing compute-aware evaluations reduce heterogeneous resources into FLOPs, which is difficult to estimate for black-box models and fails to capture resource-specific constraints. To provide a comparable evaluation basis, we introduce Fair-ASR, an evaluation protocol for black-box jailbreak attacks under shared target-call budgets B, using target calls as a directly observable and method-agnostic comparison axis while tracking attacker calls separately for efficiency analysis. We re-evaluate 11 representative attacks under the Fair-ASR protocol and find that attack rankings change substantially across target-call budgets, simple stochastic perturbations and hand-crafted templates remain highly competitive under equal target access, and no evaluated LLM-driven method is efficient in both target and attacker calls. Motivated by this efficiency gap, we introduce ReCode, a compositional budget-efficient attack that combines desensitization rewriting with two effective low-cost primitives identified by Fair-ASR. Under a budget of 20 target calls, ReCode achieves 85% ASR on GPT-5 while requiring only 7.19 attacker calls per request on average, showing strong efficiency in both target and attacker calls.

Comments29 pages, 8 figures, 13 tables. Code: https://github.com/xsddys/Fair-ASR

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑