arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向书写系统级别的字节级BPE分词器适配

Writing-System-Level Tokenizer Adaptation for Byte-Level BPE

Bohdan Didenko

arXiv 2608.00582首次发表:更新:

发表机构

Lviv Polytechnic National University(利沃夫理工国立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对字节级BPE分词器在低资源语言上效率低的问题,提出BPE引导插入等方法,在乌克兰语适配中显著降低token计数,同时保持英语等语言的token计数变化极小,且保留了大部分相同ID的模型-词汇表行。

AI 中文摘要

预训练的字节级BPE分词器对低资源语言的分词效率较低。替换分词器会改变几乎所有token ID的含义,而词汇扩展会增大模型的嵌入层和输出矩阵。我们研究事后适配,该方法保持模型-词汇表大小固定,并将保留大部分现有token到ID的分配作为构建时兼容性属性。直接从特定语言的分词器迁移token无法保证通过目标BPE合并图推导得出:插入的条目可能与目标的贪心合并排名冲突。我们将此失败形式化为合并排序问题,并提出BPE引导插入方法,该方法通过目标可达的分解构建每个迁移的token。我们的流程使用感知书写系统的行选择来限制附带碎片化,重建目标书写系统的字节级前提条件,并应用引导插入以保持合并图可达性。在Nemotron和GPT-OSS的乌克兰语适配中,该方法将token计数分别降低33.5%和36.6%,对英语及所评估的四种语言欧洲聚合体的变化保持在0.05%以内,且在相同ID下保留了78.5%/77.3%的原始模型-词汇表行。约束匹配的全局方法和基于频率的移除方法实现了类似的乌克兰语压缩,但使英语/欧洲token计数增加了0.7-2.2%;全新的相同规模再训练对乌克兰语的压缩效果略好,但几乎未保留任何相同ID的行,且使英语token计数增加了7.6-8.6%。重新分配使所评估的三种语言西里尔字母微聚合体的token计数增加了6.7%/10.1%。结构审计发现,28134/45398个插入的BPE节点在常规排名合并下均可达,且无保留的相同ID模型-词汇表条目被新破坏。我们发布所有分词器和代码。

英文摘要

Pretrained byte-level BPE tokenizers can segment underrepresented languages inefficiently. Replacing a tokenizer changes the meaning of nearly every token ID, while vocabulary expansion enlarges the model's embedding and output matrices. We study post-hoc adaptation that keeps the model-vocabulary size fixed and preserves most existing token-to-ID assignments as a construction-time compatibility property. Directly transferring tokens from a language-specific tokenizer does not guarantee derivability through the target BPE merge graph: an inserted entry can conflict with the target's greedy merge ranks. We formalize this failure as the merge ordering problem and introduce BPE-guided insertion, which builds each transferred token through a target-reachable decomposition. Our pipeline uses script-aware row selection to limit collateral fragmentation, reconstructs target-script byte-level prerequisites, and applies guided insertion to maintain merge-graph reachability. On Ukrainian adaptations of Nemotron and GPT-OSS, it reduces token counts by 33.5% and 36.6%, keeps changes on English and the evaluated four-language European aggregate within 0.05%, and retains 78.5%/77.3% of original model-vocabulary rows at the same IDs. Constraint-matched global and frequency-based removal achieve similar Ukrainian compression but increase English/European token counts by 0.7-2.2%; fresh same-size retraining compresses Ukrainian slightly more but retains effectively no same-ID rows and increases English token counts by 7.6-8.6%. The reallocation increases token counts on the evaluated three-language Cyrillic micro-aggregate by 6.7%/10.1%. Structural audits find all 28,134/45,398 inserted BPE nodes reachable under ordinary rank-ordered merging and no retained same-ID model-vocabulary entry newly broken. We release all tokenizers and code.

Comments15 pages. Accepted for poster presentation at the Second Tokenization Workshop (TokShop) at COLM 2026 (non-archival)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑