arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21088cs.CR

起源即一切:面向结构信任边界分离的溯源感知Transformer

Origin Is All You Need: Provenance-Aware Transformers for Structural Trust-Boundary Separation

Yuxuan Zhang, Jeff Huang, Guofei Gu

首次发表
浏览论文内容

中文总结 AI 辅助

针对间接提示注入,提出溯源感知Transformer,通过来源嵌入和注意力偏置实现权威与非权威来源的结构分离,在保持实用性的同时稳健抵抗攻击。

中文摘要 AI 辅助

间接提示注入(IPI)仍然是大型语言模型(LLM)系统面临的核心安全与保障挑战,因为标准Transformer在架构上缺乏对来源权威性的概念。检索到的文档、用户输入和系统指令都通过相同的无差别注意力机制处理,迫使模型仅凭措辞推断哪些内容应被遵从、哪些应被视为数据。我们提出了溯源感知Transformer(Provenance-Aware Transformers),这是一种溯源感知防御方法,使应用提供的来源标签在模型内部变得可操作。每个输入令牌被分配一个编码其来源的环ID,模型通过来源嵌入、可学习的来源注意力偏置以及可学习的来源缩放因子进行增强,从而在归一化下保留溯源信息。由此产生的架构在生成过程中强制执行权威来源与非权威来源之间的结构边界。为了在已发布的预训练模型上实例化该架构,我们提出了一种两阶段微调流程,在环约束下教会模型来源语义和任务行为。评估表明,溯源感知Transformer在分布内和分布外场景下均能保持对IPI的稳健抵抗,同时保持与基础预训练模型相当的实用性。更广泛地说,我们的工作表明,将溯源作为一等架构信号暴露出来,可以将LLM安全对齐从脆弱的模式匹配转向明确的信任分离。

英文摘要

Indirect prompt injection (IPI) remains a central safety and security challenge for large language model (LLM) systems because standard transformers lack architectural notion of source authority. Retrieved documents, user inputs, and system instructions are all processed through the same undifferentiated attention mechanism, forcing the model to infer from wording alone what should be obeyed and what should be treated as data. We propose Provenance-Aware Transformers, a provenance-aware defense that makes application-supplied source labels actionable inside the model. Each input token is assigned a ring ID encoding its origin, and the model is augmented with origin embeddings, a learnable origin attention bias, and a learnable origin scale that preserves provenance under normalization. The resulting architecture enforces a structural boundary between authoritative and non-authoritative sources during generation. To instantiate this architecture on released pretrained models, we propose a two-stage fine-tuning pipeline to teach the model origin semantics and task behavior under ring constraints. Evaluation shows that Provenance-Aware Transformers maintain robust resistance to IPI both in-distribution and out-of-distribution while preserving utility comparable to the base pretrained model. More broadly, our work shows that exposing provenance as a first-class architectural signal can shift LLM safety alignment from brittle pattern matching toward explicit trust separation.

发表机构

  • Texas A&M University(德克萨斯农工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑