arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLM 在何处决定打破规则?提示注入遵从性的机制定位

Where Do LLMs Decide to Break the Rules? Mechanistic Localization of Prompt Injection Compliance

Rui Wen, Jiayang Liu, Zeyu Yang, Jun Sakuma, Lu Sun

arXiv 2609.37737首次发表:更新:

发表机构

RIKEN AIP; Institute of Science Tokyo; Nanyang Technological University; Singapore University of Technology and Design; Tohoku University(RIKEN先进智能研究中心; 东京科学大学; 南洋理工大学; 新加坡科技设计大学; 东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过因果激活修补定位LLM遵从提示注入攻击的决策位置,发现晚期瓶颈层是关键,修补可逆转77-92%的遵从,且该层也是最优检测点。

AI 中文摘要

当提示注入攻击成功时,大型语言模型(LLM)会放弃其被分配的系统角色,转而遵从对抗性指令。虽然先前的工作已广泛量化了这种情况发生的频率,但我们提出了一个更根本的问题:在网络的哪个内部位置,模型实际上做出了打破规则的决定?通过对五个模型(参数规模从 4B 到 32B)进行逐层因果激活修补,我们发现了一个清晰的分离现象:攻击信息从第一层起即可被线性解码,然而对模型行为的因果影响力在网络的最后三分之一处的一个晚期瓶颈层之前可以忽略不计。修补该瓶颈层可在 77% 至 92% 的案例中逆转遵从行为。我们表明,遵从机制占据一个紧凑的线性子空间(在 4B 和 14B 模型中秩为 8,在 32B 时扩展到秩 64),并且在不同模型家族中具有架构稳定性。最后,我们通过证明该因果峰值层也是检测攻击的表征最优位置来验证我们的机制解释,其性能优于在诸如 leetspeak 替换等表面混淆下性能下降的早期层分类器。因果影响力与检测性能之间的这种一致性提供了汇聚证据,表明晚期瓶颈层捕获了与决策相关的计算,而不仅仅是干预伪影的反映。

英文摘要

When a prompt injection attack succeeds, a Large Language Model (LLM) abandons its assigned system role to comply with an adversarial instruction. While prior work has extensively quantified how often this occurs, we ask a more fundamental question: where inside the network does the model actually decide to break the rules? Using layer-by-layer causal activation patching across five models (4B to 32B parameters), we find a clear dissociation: attack information is linearly decodable from the first layer, yet causal leverage over the model's behavior is negligible until a late-layer bottleneck in the final third of the network. Patching this bottleneck reverses compliance in 77--92\% of cases. We show that the compliance mechanism occupies a compact linear subspace (rank-8 in 4B and 14B models, scaling to rank-64 at 32B) and is architecturally stable across varying model families. Finally, we validate our mechanistic account by showing that this causal peak layer is also the representationally optimal site for detecting attacks, outperforming early-layer classifiers that degrade under surface-level obfuscation such as leetspeak substitution. This alignment between causal leverage and detection performance provides converging evidence that the late-layer bottleneck captures decision-relevant computation rather than merely reflecting an artifact of the intervention.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑