智能体的决策依据是什么?通过定位行为引导指令来裁决未授权行为
What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions
浏览论文内容
中文总结 AI 辅助
针对LLM智能体易受动态注入攻击的问题,提出Attnlocate框架,通过目标检测定位行为引导指令以裁决未授权行为,在多场景下表现优异且可跨模型迁移。
中文摘要 AI 辅助
集成了外部资源的大语言模型(LLM)智能体具备复杂任务处理能力,但统一的自然语言上下文通道使其易受注入攻击:不可信的外部数据可能在LLM推理过程中被动态解析为行为引导指令,进而破坏智能体的决策。现有防御措施聚焦于输入/输出层面的恶意内容静态检测或隔离,对于模型推理过程中出现的这类动态诱导行为检测不足。本文提出Attnlocate,这是一个用于细粒度定位真正影响工具调用决策的上下文片段(即行为引导指令)的运行时框架。Attnlocate将该定位问题转化为目标检测任务,旨在检测注意力矩阵中由行为引导指令引发的独特激活轨迹。具体而言,我们设计了一种多头、多层注意力聚合方案,以构建适配目标检测的词元级特征空间;随后部署配备无锚点检测头的1-D U-Net来检测这些片段;最后,基于检测到的行为引导片段来源提供者的权限,Attnlocate动态裁决恶意调用尝试。我们在来自五个LLM家族的十种智能体配置上评估了Attnlocate,涵盖间接提示注入和工具中毒场景。Attnlocate的平均交并比(IoU)为0.743,平均受试者工作特征曲线下面积(AUROC)为0.956,在0.067的假阳性率下达到0.934的真阳性率。它还能有效跨未见模型迁移,且支持权限策略适配而无需重新训练。
英文摘要
LLM agents integrated with external resources gain complex task capabilities, yet the unified natural-language context channel makes them vulnerable to injection attacks: untrusted external data may be dynamically parsed as behavior-guiding instructions during LLM inference, thereby subverting the agent's decision. Existing defenses focus on static detection or isolation of malicious content at the input/output level, remains insufficient for detecting such dynamic inducements that arise during model reasoning. We propose Attnlocate, a runtime framework for fine-grained localization of context spans that genuinely influence tool-calling decisions, i.e., behavior-guiding instructions. Attnlocate casts this localization problem as an object detection task, aiming to detect the distinctive activation traces induced by behavior-guiding instructions within the attention matrix. Specifically, we design a multi-head, multi-layer attention aggregation scheme to construct a token-level feature space tailored for object detection. Then, a 1-D U-Net equipped with an anchor-free detection head is deployed to detect these spans. Finally, based on the authority of the provider from which the detected behavior-guiding spans originate, Attnlocate dynamically adjudicates malicious invocation attempts. We evaluate Attnlocate across ten agent configurations from five LLM families, covering scenarios involving indirect prompt injection and tool poisoning. Attnlocate achieves a mean IoU of 0.743, an average AUROC of 0.956, and a 0.934 true-positive rate at 0.067 false-positive rate. It also transfers effectively across unseen models and supports authority policy adaptation without retraining.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- University of Washington(华盛顿大学)
- Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。