arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EvoSafeHarness:为智能体安全演化模型特定与领域特定的防护机制

EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents

Nanxi Li, Yingzi Ma, Yulong Cao, Edward Suh, Bo Li, Dawn Song, Chaowei Xiao

arXiv 2609.05903首次发表:更新:

发表机构

Johns Hopkins University; University of Wisconsin–Madison; NVIDIA; University of Illinois Urbana–Champaign; UC Berkeley(约翰斯·霍普金斯大学; 威斯康星大学麦迪逊分校; 英伟达; 伊利诺伊大学厄巴纳-香槟分校; 加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EvoSafeHarness通过联合搜索自然语言策略和代码逻辑,为冻结模型在目标领域合成安全防护机制,在多个基准上显著提升安全-效用权衡,优于固定专家设计。

AI 中文摘要

大型语言模型(LLM)智能体正在将语言转化为现实世界的影响,这使得针对间接提示注入和直接有害请求的安全防护变得至关重要。系统级安全防护机制在模型级防御之外增加了一个执行层,但现有的防护机制通常由专家一次性设计,并应用于异构的模型和领域。有效的保护依赖于部署环境:不同模型在效用下降前所需的执行程度各不相同,而不同领域在需要治理的效果、状态和动作序列上也存在差异。对一个模型足够严格的防护机制可能会过度阻止另一个模型,而跨领域迁移的策略可能会遗漏应用特定的安全关系。我们提出了EvoSafeHarness,一个安全特定的优化框架,为冻结模型在目标领域中合成可部署的防护机制。它联合搜索自然语言策略和可执行代码逻辑,由模型行为、领域规范和新鲜上下文对抗性审查引导,以拒绝基准特定的规则。在四个智能体基准系列中,EvoSafeHarness实现了比固定专家设计的防御更强的安全-效用前沿。在DecodingTrust-Agent上,它以3.3个百分点的效用成本将平均攻击成功率从45.6%降至10.0%,并在15个单元格中的14个中取得了最佳得分。在AgentDojo上,它在0.0%攻击成功率下达到82.8%的效用,是CaMeL在同一操作点上效用的两倍,并且无需修改即可迁移到未见过的AgentDyn套件。它还在Agent-SafetyBench上为每个受害者取得了最佳得分,并在16次细化预算的自适应PAIR攻击下将平均攻击成功率保持在20%以下。分析表明,领域语义决定了需要哪些安全关系和轨迹状态,而模型和运行时行为决定了这些关系应在何处以及如何执行。

英文摘要

Large Language Model (LLM) agents are turning language into real-world effects, making safety necessary against both indirect prompt injections and direct harmful requests. System-level safety harnesses add an enforcement layer beyond model-level defenses, but existing harnesses are usually designed once by experts and applied across heterogeneous models and domains. Effective protection is deployment-dependent: models differ in how much enforcement they need before utility declines, while domains differ in the effects, state, and action sequences that must be governed. A harness that is strict enough for one model may over-block another, and a policy that transfers across domains may miss application-specific safety relations. We present EvoSafeHarness, a safety-specific optimization framework that synthesizes a deployable harness for a frozen model in a target domain. It jointly searches a natural-language policy and executable code logic, guided by model behavior, domain specifications, and fresh-context adversarial review to reject benchmark-specific rules. Across four agent benchmark families, EvoSafeHarness achieves a stronger safety-utility frontier than fixed expert-designed defenses. On DecodingTrust-Agent, it reduces average attack success rate from 45.6% to 10.0% at a 3.3-point utility cost and achieves the best score in 14 of 15 cells. On AgentDojo, it reaches 82.8% utility at 0.0% ASR, twice CaMeL's utility at the same operating point, and transfers unchanged to unseen AgentDyn suites. It also achieves the best score on Agent-SafetyBench for every victim and keeps mean ASR below 20% under adaptive PAIR attacks with a refinement budget of 16. Analysis shows that domain semantics determine which safety relations and trajectory state are needed, while model and runtime behavior determine how and where those relations should be enforced.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑