arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

威胁保持表示敏感性在智能体安全基准中的应用

Threat-Preserving Representation Sensitivity in Agent-Security Benchmarks

Neeraj Karamchandani, Piyush Nagasubramaniam, Xinhong Xie, Sencun Zhu, Dinghao Wu

arXiv 2610.03585首次发表:更新:

发表机构

The Pennsylvania State University(宾夕法尼亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出威胁保持表示敏感性(TPRS)指标,证明智能体安全基准中表示变化显著影响攻击成功率,主张鲁棒性评估需基于多种表示而非单一分数。

AI 中文摘要

针对基于大语言模型的智能体的安全基准通常将攻击成功率(ASR)作为衡量模型鲁棒性的指标,并使用这些分数来比较不同的模型和防御机制,假设它们描述了智能体的安全性。在本文中,我们探讨了这种表示是否也会影响基准的测量。为了衡量基准表示的影响,我们引入了威胁保持表示敏感性(TPRS),它衡量在保持底层任务、有害行为、安全策略、真实情况、环境和评估标准不变的情况下,改变智能体可见表示时ASR的变化程度。在Agent Security Bench(ASB)上,将威胁相关的工具名称替换为威胁中立的名称,使得GPT-5-mini的已提交攻击成功率提高了11.67个百分点,Claude Haiku 4.5提高了13.21个百分点。在MCPTox上,将原始的中立工具名称替换为明确的威胁相关名称,使得GPT-5-mini的ASR降低了11.00个百分点,Claude Haiku 4.5降低了4.11个百分点。在AgentDojo上,向攻击相关工具添加威胁相关措辞仅使GPT-4o-mini的ASR改变了0.50个百分点,但在需要该工具的任务上,良性效用下降了5.36个百分点。我们在MCPTox上进行了一项实验,观察到在令牌数量、长度和大小写上匹配的威胁中立名称重现了威胁明确名称产生的大部分变化(在GPT-5-mini上为11.00个百分点中的8.54个百分点)。结果表明,在一种表示下测得的安全分数可能无法推广到同一安全问题的威胁保持表示。因此,鲁棒性声明应通过在一组受控的威胁保持表示上的性能来支持,而不是依赖于单一的表示相关分数。

英文摘要

Security benchmarks for LLM-based agents often report the attack success rate (ASR) as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe the security of the agent. In this paper, we explore whether it also influences the benchmark's measurement. To measure the effect of the benchmark representation, we introduce threat-preserving representation sensitivity (TPRS), which measures how much the ASR changes when we change the agent-visible representation while holding the underlying task, harmful action, security policy, ground truth, environment, and the evaluation criteria fixed. On Agent Security Bench (ASB), replacing threat-related tool names with threat-neutral names raises the committed attack success rate by 11.67 percentage points on GPT-5-mini and by 13.21 points on Claude Haiku 4.5. On MCPTox, replacing the original neutral tool name with an explicit threat-related name lowers the ASR by 11.00 percentage points on GPT-5-mini and 4.11 points on Claude Haiku 4.5. On AgentDojo, adding threat-related wording to the attack-relevant tool changes ASR by only 0.50 percentage points on GPT-4o-mini, yet the benign utility falls by 5.36 points on tasks requiring that tool. We ran an experiment on MCPTox where we observed that a threat-neutral name matched on token count, length, and casing reproduces most of the shift produced by the threat-explicit name (8.54 of 11.00 points on GPT-5-mini). The results show that a security score measured under one representation may fail to generalize across threat-preserving representations of the same security problem. Robustness claims should therefore be supported by performance across a controlled set of threat-preserving representations rather than relying on a single representation-dependent score.

Comments12 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑