arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23873cs.AIcs.CLcs.CRcs.LG

语义覆盖层:通过超越 token 和引导向量的注释缓解提示注入攻击

Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors

Joshua Penman

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出 Semantic Overlays 技术,通过在冻结语言模型残差流上应用小型学习适配器构建带外注释通道,有效缓解提示注入攻击,在多个基准测试中大幅降低攻击成功率且保持模型效用。

中文摘要 AI 辅助

语言模型所接触的一切都是 token。服务栈知道每个文本段的类型——用户输入、工具输出、指令等,但模型必须自行跟踪这些信息,且可能会丢失线索或产生混淆:文本可以被构造得看似任何内容。提示注入正是利用了这一现象的自然漏洞:攻击者通过扰乱模型对文本段身份的理解,可诱导其执行不受欢迎甚至潜在危险的操作。在模型输入中添加非文本通道(一种超越文本传递文本段身份的方式)可缓解这类攻击。因此,我们提出一种名为语义覆盖层(Semantic Overlays)的通用引导技术:在选定的预填充位置向冻结模型的残差流应用小型学习适配器。在文本段上覆盖一层覆盖层,可创建一个 token 无法复制的带外注释通道。与引导向量不同,Semantic Overlays 经过训练、具有适应性且可选择性应用。覆盖层可编码复杂语义,重塑模型对标记文本段的感知:当要求在断言为不同编程语言的覆盖层下复制代码片段时,模型会忠实地将片段重写为断言的语言。覆盖层还具有可组合性,允许透明读取底层内容,并可携带复杂有效载荷,包括模型将遵循的指令。将文本段标记为“不可执行”的覆盖层可防御在不可信上下文中添加指令的广泛提示注入攻击。我们在提示注入基准测试中报告了优异结果:SEP 分离率从 24.3% 提升至 96.5% 且效用不变(采用我们的评分规则;我们还修正了已发布评分器的缺陷),TensorTrust 攻击成功率从 34.8% 降至 6.6%,所有四类 PIArena 攻击族的合规率均降至 0%,同时标记文本段保持可读(精确复制率为 92.5%)。

英文摘要

Everything a language model sees is tokens. The serving stack knows what each span is -- user input, tool output, instructions -- but the model must keep track of that itself, and can lose track or be confused: text can be written to read like anything. Prompt injection is a natural exploit of this phenomenon. By scrambling the model's understanding of span identity, an attacker can induce unwanted and dangerous actions. Adding a non-textual channel to the model's input -- a way to communicate span identity beyond text -- mitigates this class of attack. We thus introduce a general steering technique called Semantic Overlays: small learned adapters applied at chosen prefill positions to a frozen model's residual stream. Laying an overlay over a span creates an out-of-band annotation channel that cannot be replicated by tokens. Unlike steering vectors, Semantic Overlays are trained, adaptable, and selectively applied. An overlay can encode complex semantics that reshape how the model perceives the marked span: asked to copy a code snippet under an overlay asserting a different programming language, the model rewrites the snippet in the asserted language. Overlays compose, allow transparent reading of underlying content, and can carry complex payloads -- including imperatives the model will follow. An overlay which marks a span as "non-executable" defends against the broad class of prompt injections that add instructions in untrusted context. We report strong results on five prompt injection benchmarks: SEP separation rises from 24.3% to 99.0% with utility unchanged (our scoring rule; we correct a defect in the published grader), TensorTrust attack success falls from 34.8% to 6.2%, AlpacaFarm from 99.0% to 0%, and the overlay beats every published PIArena defense that leaves the model able to answer -- while marked spans stay readable, all at >95% character similarity to the original.

补充信息

↑