发表机构
EPFL(洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出TokEval框架,通过信息论、结构敏感等指标评估分词器,经实验验证其可预测下游模型性能,助力更原则性的分词器评估。
AI 中文摘要
语言模型分词器的选择通常仅经过极少的评估,尽管其设计选择会直接影响模型性能,部分原因在于人们对哪些分词器属性会影响下游性能的哪些方面认识有限。我们推出了TokEval,这是一个分词器评估指标框架,它超越了生育力、压缩率等标准指标,以捕捉具有语言学和结构意义的属性,例如UTF-8字符边界完整性、数学中的数字位值边界对齐。为验证这些指标是否能预测下游模型性能,我们开展了受控语言模型预训练实验,仅改变分词器的训练数据混合比例、预分词策略和训练算法。我们在每字节比特数(一种与分词器无关的困惑度版本)及多个基准上评估所得模型,这些基准涵盖语言理解、数学推理和代码生成。实验表明,不同的内在属性对模型能力有不同影响:信息论指标可预测语言建模能力(斯皮尔曼相关系数最高达0.80),而结构敏感指标(如测量数字和换行处理的指标)与任务准确率相关。我们希望TokEval能实现更原则性的分词器评估,在两者一致的地方用内在测量替代预训练搜索。
英文摘要
Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstream performance. We introduce TokEval, a framework of tokenizer evaluation metrics that goes beyond standard measures like fertility and compression rate to capture linguistically and structurally meaningful properties, e.g., UTF-8 character boundary integrity and digit place-value boundary alignment for mathematics. To validate whether these metrics are predictive of downstream model performance, we conduct controlled language model pretraining experiments, varying solely the tokenizers' training data mixture, pretokenization strategy, and training algorithm. We evaluate the resulting models on bits-per-byte (a tokenizer-agnostic version of perplexity) and several benchmarks, spanning linguistic understanding, mathematical reasoning, and code generation. Our experiments suggest that different intrinsic properties have different impacts on model abilities: information-theoretic metrics predict language modeling abilities (Spearman rho up to 0.80), while structure-sensitive metrics, such as those measuring digit and line-break handling, correlate with task accuracy. We hope TokEval enables more principled tokenizer evaluation, replacing pretraining sweeps with intrinsic measurement wherever the two agree.
CommentsPublished as a conference paper at COLM 2026; Library hosted at https://github.com/cimeister/tokenizer-intrinsic-evals