发表机构
Florida International University(佛罗里达国际大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Pruned BPE方法,通过训练后可见性剪枝与令牌重新分配提升BPE词汇效率,在不增加模型可见词汇规模的情况下减少编码长度,取得了优于标准BPE的效果。
AI 中文摘要
字节对编码(BPE)被广泛用于子词分词,但标准BPE会将所有学习到的合并令牌暴露给下游模型,包括那些主要作为中间构建单元、在最终编码语料中极少出现的令牌。本文提出Pruned BPE,一种将合并构建与模型可见词汇选择分离的训练后可见性剪枝及令牌重新分配方法。在标准BPE训练后,通过最终曝光度评估令牌;低曝光令牌被保留为仅内部使用的合并节点,而它们的可见词汇槽则被重新分配给通过恢复训练学习到的更高曝光候选。编码时,仅内部使用的令牌会递归扩展为可见后代,同时保留原始BPE合并顺序。在两个不重叠的以英语和汉语为主的语料及其组合上开展的实验表明,在相同训练语料、评估语料和模型可见词汇规模下,Pruned BPE能持续减少编码长度;在40%曝光阈值下,同语料评估时的减少幅度约为0.27%至0.36%。在使用共享精确最小令牌动态规划编码器的仅词汇评估中,Pruned BPE仍保持约0.23%至0.31%的优势,说明改进源于更高效的可见词汇。这些收益占边际减少量的相当比例,否则需再添加2000个标准BPE令牌才能实现约1.5%至3.8%的边际减少。定性分析显示,仅内部使用的令牌包含可复用的英语片段、汉语组件、部分UTF-8字节序列及结构化文本片段。结果表明,训练后可见性剪枝可在不增加语言模型可见词汇的前提下提升BPE词汇效率。
英文摘要
Byte Pair Encoding (BPE) is widely used for subword tokenization, but standard BPE exposes every learned merge token to the downstream model, including tokens that mainly serve as intermediate construction units and rarely appear in the final encoded corpus. This paper proposes Pruned BPE, a post-training visibility-pruning and token-reallocation method that separates merge construction from model-visible vocabulary selection. After standard BPE training, tokens are evaluated by final exposure. Low-exposure tokens are retained as internal-only merge nodes, while their visible vocabulary slots are reassigned to better-exposed candidates learned through resumed training. During encoding, internal-only tokens are recursively expanded into visible descendants while the original BPE merge order is preserved. Experiments on two non-overlapping English- and Chinese-dominated corpora and their combination show that Pruned BPE consistently reduces encoded length relative to Standard BPE at the same training corpus, evaluation corpus, and model-visible vocabulary size. At a 40% exposure threshold, the reduction is approximately 0.27%--0.36% on same-corpus evaluations. In a vocabulary-only evaluation using a shared exact minimum-token dynamic-programming encoder, Pruned BPE retains an advantage of approximately 0.23%--0.31%, indicating that the improvement arises from a more efficient visible vocabulary. These gains represent a meaningful fraction of the approximately 1.5%--3.8% marginal reduction that would otherwise require adding another 2K Standard BPE tokens. Qualitative analysis shows that internal-only tokens include reusable English fragments, Chinese components, partial UTF-8 byte sequences, and structured-text fragments. The results indicate that post-training visibility pruning can improve BPE vocabulary efficiency without increasing the vocabulary exposed to the language model.
Comments18 pages, 2 figures, 4 tables, and 1 algorithm