arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你的智能体大模型会秘密编码间接提示注入暴露的潜在信号

Your Agentic LLMs Secretly Encode Indirect Prompt-Injection Exposure in Hidden States

Jianshuo Dong, Yiming Liu, Maosen Zhang, Nan Deng, Peng Xu, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu

arXiv 2608.02657首次发表:更新:

发表机构

Tsinghua University; MatrixOrigin; Nanyang Technological University; SiliconProspect AI(清华大学; 矩阵起源; 南洋理工大学; 硅景人工智能)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究智能体大模型的间接提示注入暴露问题,通过探测、防御、解释三方面分析,提出AGRI防御方法,可大幅降低攻击成功率且保持任务效用

AI 中文摘要

智能体大模型(Agentic LLMs)易受间接提示注入(IPI)攻击,例如外部工具结果中隐藏的恶意侧任务。尽管已有诸多应对此类威胁的研究,但当智能体大模型暴露于IPI攻击(我们称之为IPI暴露)时,其内部机制尚不明确。本文从三个方面深入研究该问题:(1)探测:针对包括7530亿参数的GLM-5.2在内的6种模型,基于生成前隐藏状态训练的简单线性探测器可预测LLM的IPI暴露,这些探测器在未见过的攻击、智能体指令和任务套件上实现90%以上的AUROC,且在自适应攻击和跨语言场景下表现出高鲁棒性;(2)防御:我们的思维链(CoT)测量揭示了识别-行动差距:尽管模型会编码此类信号,但常无法将其转化为安全行动。随后我们提出AGRI,一种由探测器门控的基于推理的防御方法,可按需前置抗注入推理。在复杂的AgentDojo设置中,AGRI大幅降低了攻击成功率,例如在Qwen3.5-27B上从34.6%降至0%,同时基本保持干净任务的效用;(3)解释:我们提出一种分析框架,可识别与探测器捕获信号最相关的自然语言解释,生成的特征在不同模型间存在差异:潜在信号可与直接IPI暴露声明或间接操作线索对齐。代码可获取:this https URL

英文摘要

Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While many efforts have sought to address this threat, little is known about the internals of agentic LLMs when they are exposed to IPI attacks. For simplicity, we refer to this condition as IPI exposure. In this paper, we study IPI exposure from three perspectives. (1) Probing: Across eight models, including the 753B-parameter GLM-5.2 and the 2.8T-parameter Kimi-K3, simple linear probes trained on pre-generation hidden states can predict LLMs' IPI exposure. These probes achieve 0.90+ AUROC on unseen attacks, agent instructions, and task suites; they remain robustly predictive under adaptive attacks and in cross-lingual settings. (2) Defense: We reveal and diagnose a knowledge-action gap: post-trained LLMs encode signals predictive of IPI exposure, yet do not reliably bind these signals to safe agentic actions. We therefore introduce a probe-gated reasoning-based defense to bridge this gap at test time. On difficult AgentDojo settings, it substantially reduces attack success rate, e.g., from 34.6% to 0% on Qwen3.5-27B, and better preserves clean-task utility than the baselines. (3) Explanation: We introduce an analysis framework that identifies natural-language explanations strongly correlated with probe-captured signals. The resulting profiles differ across models: latent signals can align with either direct IPI-exposure sensing or indirect operational cues. Code is available at https://github.com/jianshuod/IPI-exposure-signal.

Commentsv2. Compared to v1, we include the Kimi-K3 model's results and conduct a series of new experiments on understanding probe generalization

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑