发表机构
Beijing University of Posts and Telecommunications; Beihang University; Tsinghua University(北京邮电大学; 北京航空航天大学; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究发现模式格式工具规格是AI智能体安全下降的主因,提出SafeKeep防护措施,可提升有害请求弃权率、降低攻击成功率,性能优于现有防护且保留任务处理能力。
AI 中文摘要
AI智能体通过外部工具扩展了大型语言模型(LLMs),使其能够执行复杂任务并将模型输出转化为有重要影响的现实世界行动。然而,LLMs作为智能体部署时,安全性往往大幅下降,这种下降的根源仍知之甚少。在本文中,我们将模式格式的工具规格确定为智能体安全性下降的主要来源,并通过白盒表示分析表明,它们会削弱模型的内部拒绝信号,导致不安全的工具执行。基于这一发现,我们提出了SafeKeep,一种推理时的安全防护措施,它将安全判断与工具执行解耦:使用扁平化文本工具规格评估请求,同时保留原始模式格式的规格用于执行。在两个代表性基准和四个LLMs(包括白盒和黑盒模型)上,SafeKeep将有害请求的平均弃权(不执行)率从23.8%提升至70.6%,并将观测级提示注入下的平均攻击成功率从25.6%降至2.5%。它还优于现有安全防护措施,且保留了任务处理能力。我们在该https URL发布代码和数据。
英文摘要
AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential real-world actions. Yet LLMs often become substantially less safe when deployed as agents, and the source of this degradation remains poorly understood. In this paper, we identify schema-formatted tool specifications as a primary source of agent safety degradation and show, through white-box representation analysis, that they weaken the model's internal refusal signals and contribute to unsafe tool execution. Building on this finding, we propose SafeKeep, an inference-time safeguard that decouples safety judgment from tool execution: it assesses requests using flattened textual tool specifications while retaining the original schema-formatted specifications for execution. Across two representative benchmarks and four LLMs, including both white-box and black-box models, SafeKeep increases the average refusal rate for harmful requests from 23.8% to 70.6% and reduces the average attack success rate under observation-level prompt injection from 25.6% to 2.5%. It also outperforms existing safeguards and preserves task-handling capability. We release the code and data at https://github.com/snowcatsmoking/SafeKeep .