发表机构
Zhejiang University; Binjiang Institute of Zhejiang University; Hong Kong Baptist University(浙江大学; 浙江大学滨江研究院; 香港浸会大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过因果追踪揭示间接提示注入中角色信号的可读性与行为干预有效性之间的差异,发现深度、位置和上下文对干预效果有重要影响。
AI 中文摘要
间接提示注入导致LLM智能体遵循外部数据中嵌入的指令。一个探针可能区分指令与数据,而无需识别改变下一步行动的状态编辑。我们通过反事实角色探针、组件级激活修补以及对AgentDojo轨迹的单独干预来研究这一差距。角色解码在内容和格式变化中得以存续。在受控的Qwen测试中,它先于沿独立估计的角色方向进行的修补所产生的强工具选择效应。在AgentDojo上,从被劫持和抵抗的训练轨迹中估计的方向在行动前和注入跨度位置降低了攻击成功率,但在随机位置效果甚微。在更长的Qwen轨迹中,单位置编辑在较深层变得效果较差;跨度广泛和重复编辑在同一评估集上降低了攻击成功率。移除学习到的通道子空间保持了角色解码,然而有效的干预方向在测试的通道间转移性差。这些发现区分了可读的角色信号与有效的行为干预:深度在受控工具选择中至关重要,而位置和上下文在攻击轨迹中同样重要。
英文摘要
Indirect prompt injection causes LLM agents to follow commands embedded in external data. A probe may distinguish instructions from data without identifying a state edit that changes the next action. We study this gap through counterfactual role probes, component-wise activation patching, and separate interventions on AgentDojo trajectories. Role decoding survives changes in content and format. In controlled Qwen tests, it precedes strong tool-choice effects from patches along an independently estimated role direction. On AgentDojo, directions estimated from hijacked and resisted training trajectories reduce attack success at pre-action and injected-span positions, but have little effect at random positions. In longer Qwen trajectories, single-position edits become less effective at later layers; span-wide and repeated edits reduce attack success on the same evaluation set. Removing the learned channel subspace preserves role decoding, yet effective intervention directions transfer poorly across the tested channels. These findings distinguish a readable role signal from an effective behavioral intervention: depth matters in controlled tool choice, while position and context also matter in attack trajectories.
Comments25 pages, 9 figures, including appendices