HASTE:利用稀疏证据演化智能体防护机制以应对新兴攻击
HASTE: Evolving Agent Harnesses Against Emerging Attacks Using Sparse Evidence
浏览论文内容
中文总结 AI 辅助
HASTE提出多智能体对抗演化框架,从稀疏威胁证据中自动更新智能体防护机制,以应对新兴攻击,实验表明在保持任务效用的同时降低攻击成功率。
中文摘要 AI 辅助
智能体防护机制(Agent harnesses)通过强制执行安全约束来防止不安全行为,在防御中发挥着关键作用。然而,快速涌现的攻击使人工防护机制调整难以跟上,这推动了自动化防护机制演化的需求。然而,可用于防护机制演化的信号往往稀疏,例如威胁报告和预印本中的简短描述或少量攻击示例。为解决这一局限,我们提出了HASTE,一个多智能体框架,通过安全规范生成与攻击案例生成之间的对抗性互动,从稀疏威胁证据中演化智能体防护机制。安全规范指导防护机制更新以解决已识别的安全漏洞,而攻击案例则在每次更新后探测剩余的安全漏洞。通过将评估结果反馈到这两个过程中,HASTE能够针对超出初始观察证据的新兴攻击实现防护机制演化。在多个骨干模型、攻击类型和证据形式上的实验结果表明,HASTE在保持良性任务效用的同时,持续降低了攻击成功率。代码可在以下网址获取:此https URL。
英文摘要
Agent harnesses play a critical role in defenses by enforcing safety constraints to prevent unsafe actions. However, rapidly emerging attacks outpace manual harness adaptation, motivating automated harness evolution. Yet the signals available for harness evolution are often sparse, such as brief descriptions or a few attack examples in threat reports and preprints. To address this limitation, we introduce HASTE, a multi-agent framework that evolves agent harnesses from sparse threat evidence through an adversarial interplay between safety-specification generation and attack-case generation. Safety specifications guide harness updates toward addressing identified safety vulnerabilities, while attack cases probe for remaining safety vulnerabilities after each update. By feeding evaluation outcomes back into both processes, HASTE enables harness evolution against emerging attacks beyond the initially observed evidence. Experimental results across multiple backbone models, attack types, and evidence forms show that HASTE consistently reduces attack success rates while preserving benign-task utility. The code is available at https://github.com/xxiqiao/HASTE.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- National University of Singapore(新加坡国立大学)
- Singapore Management University(新加坡管理大学)
机构由 AI 辅助整理,请以论文原文为准。