Teaching Old Tokenizers New Words: Efficient Tokenizer Adaptation for Pre-trained Models
为预训练模型教学旧分词器新词汇:高效的分词器适应方法
机构 * Institute of Computer Science, University of Tartu(塔尔图大学计算机科学研究所) ; CAIRO, Technical University of Applied Sciences Würzburg-Schweinfurt(魏玛-施维林应用技术大学)
AI总结 本文提出通过继续BPE训练扩展分词器词汇并引入基于叶节点的词汇剪枝,提升分词效率和词汇利用率,提供开放源代码工具。
Comments Accepted to Findings of EACL 2026
Journal ref Findings of the Association for Computational Linguistics: EACL 2026, pages 6492-6516