arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无名分词:针对开放权重大语言模型中控制令牌伪造的无损分词器级防御

Nameless Tokenization: A Lossless Tokenizer-Level Defense Against Control-Token Forgery in Open-Weight LLMs

Kisu Yang, Yoonna Jang, Heuiseok Lim

arXiv 2609.16984首次发表:更新:

发表机构

Korea University; Hanwha Aerospace; VAIV Company(高丽大学; 韩华航空航天; VAIV公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对开放权重LLM中控制令牌伪造问题,提出无名分词法,通过保留标识符但去除表面字符串实现无损防御,在256个分词器审计中显著提升抗伪造能力。

AI 中文摘要

开放权重语言模型会发布其聊天模板中用于标记轮次、角色和工具结果的字符串,分词器将这些字符串映射回模型所遵循的保留标识符。因此,任何能控制提示中文本的人都可以写出与服务平台栈所生成的轮次边界无法区分的边界。我们审计了256个已部署的聊天分词器。所有分词器均可被伪造,而通常建议作为修复方案的标志仍使56.6%的分词器可被伪造,因为它遗漏了智能体系统所依赖的工具和推理标记。我们提出了无名分词法,该方法为控制条目保留标识符但不提供表面字符串,因此内容编码器无法生成此类条目,消息内容得以原样到达模型。在五种分词器家族中,该方法在无攻击数据上精确复现标准令牌流,并在含分隔符文本的探针上将准确率从8.5%提升至59.9%,而净化器在此场景下失效。将分隔符的外观与其标识符分离表明,在面对裸任务指令时标识符影响甚微,但一旦系统消息告知模型将用户内容视为数据,标识符便承载了大部分伪造工具结果及大部分伪造轮次。

英文摘要

Open-weight language models publish the strings their chat templates use to mark turns, roles and tool results, which the tokenizer maps back to the reserved identifiers the model obeys. Anyone who controls text in a prompt can therefore write a turn boundary indistinguishable from one the serving stack wrote. We audit 256 deployed chat tokenizers. All are forgeable, and the flag usually recommended as a fix leaves 56.6% forgeable because it misses the tool and reasoning markers agent systems rely on. We propose nameless tokenization, which leaves the control entries with a reserved identifier and no surface string, so the content encoder cannot emit one and message content reaches the model unaltered. Across five tokenizer families it reproduces the standard token stream exactly on attack-free data and lifts accuracy on a probe of delimiter-bearing text from 8.5% to 59.9%, where sanitizers lose it. Separating a delimiter's appearance from its identifier shows the identifier matters little against a bare task instruction, but carries most of a forged tool result and most of any forged turn once the system message tells the model to treat user content as data.

Commentspreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑