AI 中文总结
本文论证攻击成功率(ASR)因六个未明确的设计选择而成为一族指标,导致跨论文比较缺乏有效性;通过259篇论文的元分析及分析研究揭示其不可比性,并提出十项报告清单以建立共享测量契约。
AI 中文摘要
攻击成功率(ASR)几乎是所有已发表的针对LLM智能体的攻击与防御评估中的首要指标。我们认为,当前使用的ASR并非单一量,而是一族由六个设计选择参数化的指标,这些选择在论文中很少被明确说明,且在整个文献中从未保持一致。我们通过两项无需专有访问权限的研究来支持这一观点。首先,对2025年2月至2026年9月间发布在arXiv上的259篇智能体安全论文进行的全文元分析发现,大多数论文既未报告其首要攻击指标的方差估计,也未报告重复运行结果:在手编码的50篇随机样本中,58%(95%置信区间44-71)如此,通过自动编码全部259篇论文,这一比例为65.3%。仅有30.9%的论文充分披露了解码相关信息,以确定其评估是否具有随机性;在我们确认使用LLM评判器的64篇论文中,29.7%报告了与人工标签的一致性检查。其次,一项分析研究表明,这些遗漏并非无关紧要:在一个包含100个实例的基准上,在常规统计功效下可检测到的ASR最小差异为18.2个百分点,而两个真实ASR相差5个百分点的防御措施,在单次运行评估中约有21%的时间被错误排序。由于六个轴中的几个会以系统依赖的方式改变ASR,由此产生的不可比性并非在比较中可抵消的恒定偏移。我们得出结论,目前跨论文的ASR比较缺乏支持,并针对我们测量到的每个失败点提出了一份包含十项的报告清单。我们的目的并非质疑任何个别结果,而是提供该领域迄今所缺少的共享测量契约。
英文摘要
Attack success rate (ASR) is the headline metric in nearly every published evaluation of attacks on, and defenses for, LLM agents. We argue that ASR as currently used is not a single quantity but a family of metrics parameterized by six design choices that papers seldom specify and never hold constant across the literature. We support this with two studies that require no proprietary access. First, a full-text meta-analysis of 259 agentic-security papers posted to arXiv between February 2025 and September 2026 finds that most report neither a variance estimate nor repeated runs for their headline attack metric: 58% (95% CI 44-71) in a hand-coded random sample of 50, 65.3% by automated coding of all 259. Only 30.9% disclose enough about decoding to establish whether their evaluation was even stochastic, and of the 64 papers we confirm use an LLM judge, 29.7% report any agreement check against human labels. Second, an analytical study shows that these omissions are not cosmetic: on a 100-instance benchmark, the minimum difference in ASR detectable at conventional power is 18.2 percentage points, and two defenses whose true ASRs differ by 5 points are ranked in the wrong order by a single-run evaluation roughly 21% of the time. Because several of the six axes shift ASR in a system-dependent way, the resulting incomparability is not a constant offset that cancels in comparison. We conclude that cross-paper ASR comparison is currently unsupported, and propose a ten-item reporting checklist targeted at each failure we measure. Our aim is not to dispute any individual result but to supply the shared measurement contract the field has so far done without.
Comments7 pages, 2 tables, 2 Figures