arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

相同字节,不同权威:聊天模板提示注入中的保留令牌表示

Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection

Yan Zhan, Yunze Song, Mengkai Hou, Wanting Zhang, Shaobo Liu, Zhijun Gao

arXiv 2609.35932首次发表:更新:

发表机构

Peking University; National University of Singapore; BYD Company Limited; Shenzhen University(北京大学; 新加坡国立大学; 比亚迪股份有限公司; 深圳大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过对比保留令牌与子词编码的伪造模板标记,测量了聊天模板提示注入中权威性的来源,发现权威性主要来自保留令牌的单一学习向量,且标准缓解措施在多数配置中无法保护工具协议令牌。

AI 中文摘要

当注入指令被包裹在模型自身的聊天模板中时,针对LLM智能体的提示注入攻击会变得强大得多。伪造的模板标记(如<|im_start|>)可以以单个保留控制令牌或一系列普通子词令牌的形式到达模型。这两种方式解码后得到完全相同的文本,并且由于分词在服务器端运行,由防御者而非攻击者决定模型接收哪一种形式。我们利用这一点来测量注入指令的权威性中有多少来自保留令牌的学习表示。将伪造标记编码为子词,同时保持文本固定并控制由此增加的额外令牌,在InjecAgent基准上,四个开放权重模型家族中的三个,攻击成功率降低了39到66个百分点,并且这一差距延续到了AgentDojo中的多轮智能体任务。在Qwen3-8B上差距为8个百分点,因为即使没有保留ID,模型仍能通过推理从文本中识别出伪造的轮次;抑制推理块将差距扩大到50个百分点。权威性位于标记位置处的单个学习向量中:标记子词向量的均值无法复现该权威性,最近普通令牌的向量在Llama-3.1上恢复了攻击,而自适应攻击者搜索非保留标记时,在四个家族中的三个找到了这样的嵌入邻居。在我们测试的每一对基础模型和指令微调模型中,指令微调增强了模型对保留标记的偏好。标准的缓解措施——一种将特殊令牌编码为普通子词的分词器选项——仅适用于配置中声明为特殊的令牌,因此在67个不同的分词器配置中的33个(覆盖Hugging Face上400个最常下载聊天模型中的255个)中,该措施无法保护智能体读取不可信工具输出所依赖的工具协议令牌,并且这一差距在该通道上持续存在。

英文摘要

Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as <|im_start|> can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The two decode to exactly the same text, and because tokenization runs on the server, the defender rather than the attacker decides which one the model receives. We use this to measure how much of the injected instruction's authority comes from the reserved token's learned representation. Encoding the forged markers as subwords, with the text held fixed and a control for the extra tokens this adds, lowers attack success on the InjecAgent benchmark by 39 to 66 percentage points on three of four open-weight families, and the gap carries over to multi-turn agent tasks in AgentDojo. On Qwen3-8B the gap is 8 points, because without reserved ids the model still recognises the forged turn from its text by reasoning; suppressing the reasoning block widens the gap to 50. The authority sits in the single learned vector at the marker position: the mean of the marker's subword vectors does not reproduce it, the vector of the nearest ordinary token restores the attack on Llama-3.1, and an adaptive attacker who searches for non-reserved markers finds such embedding neighbours on three of four families. In every base and instruction-tuned pair we test, instruction tuning strengthens the model's preference for reserved markers. The standard mitigation, a tokenizer option that encodes special tokens as ordinary subwords, applies only to tokens a configuration declares special, so in 33 of 67 distinct tokenizer configurations, covering 255 of the 400 most-downloaded chat models on Hugging Face, it leaves intact the tool-protocol tokens through which agents read untrusted tool output, and the gap persists on that channel.

Comments28 pages, 6 figures. Code: https://github.com/Byte-Authority Dataset: https://huggingface.co/datasets/YanZhanPKU/Byte-Authority-Evaluation

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑