arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GlitchPatch:通过局部重新分词修复冻结语言模型中的故障词元

GlitchPatch: Repairing Glitch Tokens in Frozen Language Models via Local Retokenization

Kunsheng Tang, Peigui Qi, Yide Song, Peijun Huang, Weiming Zhang, Nenghai Yu

arXiv 2610.04399首次发表:更新:

发表机构

University of Science and Technology of China; University of Washington; Wuhan University(中国科学技术大学; 华盛顿大学; 武汉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对冻结语言模型中的故障词元问题,提出基于局部重新分词的GlitchPatch框架,通过行为路径优化生成替换规则,在不修改模型参数的情况下实现高修复率并降低故障率。

AI 中文摘要

故障词元是词汇表中的异常条目,可能导致大型语言模型(LLM)产生与其输入不一致的输出。现有的修复方法需要访问模型内部信息,因此对于冻结检查点而言并不实用。我们研究了是否可以通过优化输入分词在模型外部修复故障词元。一项关于BPE合并规则删除的实证研究表明:(1)删除故障词元的合并规则可以修复相当一部分失败案例,但干扰共享中间合并节点的正常词元会导致整体故障率上升;(2)不同的分解粒度会产生非单调的修复率,而对正常词元的附带损害则单调增加。基于这些发现,我们提出了GlitchPatch,一种基于局部重新分词的冻结语言模型修复框架,包含两个阶段:离线阶段使用行为路径优化(BPO)为每个故障词元寻找行为上最优的替换词元序列,并将验证过的替换方案编译成规则表;在线阶段仅替换规范词元序列中匹配的故障词元ID,不修改模型参数或内部状态。在涵盖六个分词器家族的十个模型上的实验表明,GlitchPatch实现了85.10%的平均修复率,比最强基线高出14.37个百分点,并将平均故障率从14.88%降至2.27%。GlitchPatch在全词汇评估中实现了0.00%的RR,并且设计上使未匹配规则的输入保持不变。我们进一步从时间成本、语言理解和能力角度评估了修复的实际影响,支持其部署可行性。我们希望这项工作为提高分词器可靠性提供一个实用的选择。

英文摘要

Glitch tokens are anomalous vocabulary entries that can cause large language models (LLMs) to produce outputs inconsistent with their inputs. Existing repair methods require access to model internals, making them impractical for frozen checkpoints. We investigate whether glitch tokens can be repaired outside the model by optimizing the input tokenization. An empirical study on BPE merge-rule deletion reveals that (1)deleting a glitch token's merge rule can fix a substantial fraction of failures, yet disrupting normal tokens sharing intermediate merge nodes causes the overall glitch rate to rise, and (2)different decomposition granularities yield non-monotonic fix rates while collateral damage on normal tokens grows monotonically. Motivated by these findings, we propose GlitchPatch, a repair framework for frozen language models based on local retokenization, consisting of two stages: the offline stage uses Behavioral Path Optimization (BPO) to find the behaviorally optimal replacement token sequence for each glitch token and compiles validated replacements into a rule table; the online stage substitutes only the IDs of matched glitch tokens in the canonical token sequence, with no modification to model parameters or internal states. Experiments on ten models spanning six tokenizer families show that GlitchPatch achieves an 85.10% mean fix rate, outperforming the strongest baseline by 14.37 percentage points, and reduces the average glitch rate from 14.88% to 2.27%. GlitchPatch achieves a 0.00% RR in full-vocabulary evaluation and leaves rule-unmatched inputs unchanged by design. We further evaluate the practical impact of repair from the perspectives of time cost, language understanding, and capability, supporting its deployment feasibility. We hope this work provides a practical option for improving tokenizer reliability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑