发表机构
ETH Zürich; EPFL(苏黎世联邦理工学院; 洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过补全2x2设计空间,分解词元化器的优化目标与搜索过程,发现搜索过程(自底向上合并)主导性能,而非目标,为构建更优分词器提供指导。
AI 中文摘要
现代语言模型使用两种主流的词元化算法:字节对编码(BPE)和UnigramLM。这两种算法在两个正交的维度上有所不同:它们的优化目标(压缩 vs. 对数似然)和它们的搜索过程(自底向上的合并 vs. 自顶向下的剪枝)。现有的比较混淆了这些维度,使得观察到的差异究竟是源于优化了什么还是如何优化的变得不清楚。我们通过引入两种新的词元化算法来解开这两个维度,从而补全这个2x2的设计空间:BottomUpLL,一种自底向上的基于似然的词元化器,以及TopDownComp,一种自顶向下的基于压缩的词元化器。我们使用每种算法产生的词元化器训练语言模型,并改变:模型大小、词汇表大小和领域(仅英语 vs. 多语言)。在字节每比特(bits-per-byte)指标上评估模型,我们发现搜索过程——而非目标——是主导因素:自底向上的词元化器在大多数设置中一致地实现了更低的字节每比特。然而,在BLiMP任务上评估模型,显示设计选择与性能之间没有一致的关系。总的来说,我们的结果解开了词元化器设计选择对语言建模性能的影响,为其更原则性的构建提供了具体指导。
英文摘要
Two dominant tokenisation algorithms are used by modern language models: byte-pair encoding (BPE) and UnigramLM. These differ along two orthogonal axes: their optimisation objective (compression vs. log-likelihood) and their search procedure (bottom-up merging vs. top-down pruning). Existing comparisons confound these axes, making it unclear whether their observed differences stem from what is being optimised vs. how it is being optimised. We disentangle the two by introducing two new tokenisation algorithms that complete this 2x2 design space: BottomUpLL, a bottom-up likelihood-based tokeniser, and TopDownComp, a top-down compression-based tokeniser. We train language models with tokenisers produced by each algorithm, varying: model size, vocabulary sizes, and domain (English-only vs. multilingual). Evaluating models on bits-per-byte, we find that the search procedure -- not the objective -- is the dominant factor: bottom-up tokenisers consistently achieve lower bits-per-byte in most settings. Evaluating models on the BLiMP task, however, shows no consistent relationship between design choice and performance. Overall, our results disentangle the effect of tokeniser design choices on language modelling performance, offering concrete guidance for their more principled construction.
CommentsAccepted at EMNLP 2026. 20 pages, 4 figures, 10 tables. Code: https://github.com/Ahmetcanyvz/comp-vs-like