表示转变揭示多轮LLM智能体中新兴的安全风险
Representation Transitions Reveal Emerging Safety Risks in Multi-Turn LLM Agents
浏览论文内容
中文总结 AI 辅助
针对多轮LLM智能体攻击,提出DART运行时框架,通过去噪表示转变检测并归因有害行为,在六个模型上将攻击成功率从84%降至25%,优于现有防御。
中文摘要 AI 辅助
对智能体系统的多轮攻击可以将单独允许的行为组合成有害结果,这给那些孤立评估行为或状态的防御机制带来了挑战。我们表明,此类攻击会在智能体的内部表示中留下可检测的特征:有害行为表现为跨上下文更新的累积表示转变,其触发上下文可以从同一信号中识别出来。我们进一步发现,朴素的聚合会被良性的表示漂移所混淆,因为对比安全方向不需要对良性转变赋予零值。我们通过去噪方向来解决这个问题,将良性流量锚定在零值并移除其主要变化方向,且不增加运行时成本。这些发现催生了DART,一个运行时框架,用于检测和归因表示转变,并通过有针对性的提醒进行干预。在六个模型和两个多轮基准测试中,DART在MT-AgentRisk上将攻击成功率从84%降低到25%,以12%的平均误报率捕获了每一次攻击;在ASEval上,攻击成功率从97%降低到52%,良性不拒绝成本分别为8%和0%。在MT-AgentRisk上,它在所有六个模型上都优于最先进的多轮防御ToolShield:在相同协议下,ToolShield仅达到55%。去噪至关重要:在ASEval上,未去噪的监控器仅捕获7%-40%的攻击,而去噪后的监控器捕获60%-85%的攻击。同一监控器无需修改即可覆盖单轮间接注入,并且每个监控步骤仅增加0.14-0.56秒的开销,无需辅助模型,使其成为计算密集型推测性防御的轻量级补充。
英文摘要
Multi-turn attacks on agentic systems can compose individually permissible actions into harmful outcomes, challenging defenses that assess actions or states in isolation. We show that such attacks leave a detectable signature in the agent's internal representations: harmful behavior emerges as an accumulated representation transition across context updates, whose triggering context can be identified from the same signal. We further find that naive aggregation is confounded by benign representation drift, as a contrastive safety direction need not assign zero to benign transitions. We address this by denoising the direction, anchoring benign traffic at zero and removing its leading variation directions, with no runtime cost. These findings motivate DART, a runtime framework that detects and attributes representation shifts and intervenes with targeted reminders. Across six models and two multi-turn benchmarks, DART reduces attack success from 84% to 25% on MT-AgentRisk, catching every attack at a mean false-alarm rate of 12%, and from 97% to 52% on ASEval, at costs in benign non-refusal of 8% and 0%, respectively. On MT-AgentRisk, it outperforms ToolShield, the state-of-the-art multi-turn defense, on all six models: under the same protocol, ToolShield reaches only 55%. Denoising is critical: on ASEval, the undenoised monitor catches only 7%-40% of attacks, while the denoised monitor catches 60%-85%. The same monitor covers single-turn indirect injection without modification and adds only 0.14-0.56 s overhead per monitored step without requiring an auxiliary model, making it a lightweight complement to computation-heavy speculative defenses.
发表机构
- Singapore Management University(新加坡管理大学)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。