arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型分词中的计数与最小成本编码

Counting and Min-Cost Encoding for Tokenization in Large Language Models

Shuming Shi, Xiang Zhang, Hao Yu, Wenbo Fei, Changjian Wang, Zhan Wang, Guoqing Pang, Guangye Yu, Quan Lu, Ning Jiang

arXiv 2610.01127首次发表:更新:

发表机构

Mashang Consumer Finance Co., Ltd.; National-Mathematics Artificial Intelligence Institute in Chongqing (NMAII)(马上消费金融股份有限公司; 重庆国家数学人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出CNF分词器训练与MCE最小成本编码算法,通过全局优化分割和计数过滤构建词表,相比BPE在压缩率、词元效率和词表利用率上显著提升,且下游性能相当。

AI 中文摘要

主流大语言模型依赖分词器将文本编码为词元序列。对于同一文本,不同的分词器可能产生长度差异显著的词元序列。在固定模型架构下,较短的词元序列对应更低的推理时间。我们提出了一种名为计数与过滤(CNF)的分词器训练方法,以及一种名为最小成本编码(MCE)的文本编码算法。MCE 定义了一个文本片段上的成本函数,并通过全局最小化整体分割成本来确定最佳分割。CNF 通过直接计数有效子串来构建原始词表,然后在使用 MCE 分割训练语料时,基于实际词元使用情况通过过滤步骤构建最终词表。与 BPE 相比,CNF-MCE 组合提供了多项优势,包括更高的词元效率、更强的可扩展性和更低的依赖性。在六种文本类别和两种词表大小分组中,CNF-MCE 始终比所评估的 BPE 分词器实现更好的压缩效果。使用 250K 词表时,在英文网页文本上,CNF-MCE 相较于 o200k_base 和 qwen250k 分词器,压缩率分别提升了 26% 和 30%。将词表扩展到 100 万条目的英文网页文本实验表明,相对于 BPE 有持续改进,词元效率提升超过 60%,词表利用率从 52.9% 上升到 96.9%。MCE 算法不依赖于合并列表(如 BPE)或词元概率(如 UnigramLM),因此适用于广泛的词表,包括由 BPE、UnigramLM、CNF 等构建的词表。在 1.8B 和 8B 规模下从头训练的语言模型,在 11 个基准测试中取得了与使用 BPE 分词器的模型相当的平均性能。这些结果表明,CNF-MCE 可以显著提高词元效率,同时保持有竞争力的下游性能。

英文摘要

Mainstream large language models rely on a tokenizer to encode text into a token sequence. Different tokenizers may yield token sequences of substantially different lengths for the same text. With a fixed model architecture, shorter token sequences correspond to lower inference time. We propose a tokenizer training approach named Counting and Filtering (CNF) and a text encoding algorithm called Min-Cost Encoding (MCE). MCE defines a cost function over a text segment, and determines the best segmentation by globally minimizing the overall segmentation cost. CNF builds a raw vocabulary by directly counting valid substrings, and then constructs the final vocabulary through a filtering step based on actual token usage when segmenting the training corpus with MCE. The CNF-MCE conbination offers several advantages over BPE, including higher token efficiency, greater scalability, and lower dependency. Across six text categories and two vocabulary-size groups, CNF-MCE consistently achieves better compression than the evaluated BPE tokenizers. With a 250K vocabulary, CNF-MCE increases compression rate by 26% and 30% on English web text over the o200k_base and qwen250k tokenizers. Experiments scaling the vocabulary to 1M entries on English web text demonstrate sustained improvements over BPE, with a token efficiency improvement of over 60% and vocabulary utilization rising from 52.9% to 96.9%. The MCE algorithm does not depend on a merge list (as in BPE) or token probability (as in UnigramLM), making it applicable to a wide range of vocabularies, including those built from BPE, UnigramLM, CNF, and others. Language models trained from scratch at the 1.8B and 8B scales achieve comparable average performance to models using the BPE tokenizers across 11 benchmarks. These results demonstrate that CNF-MCE can improve token efficiency significantly while maintaining competitive downstream performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑