发表机构
EPFL; University of Toronto; Vector Institute(洛桑联邦理工学院; 多伦多大学; 矢量研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多语言语言模型,训练123个模型对比54种分词器,发现分词器选择对低资源语言影响更大,且可通过相关指标预测下游模型BPB排名以筛选分词器。
AI 中文摘要
分词器选择会影响多语言语言建模,但词汇容量有限,词汇大小常受约束:提升部分语言的表示往往以牺牲其他语言为代价。因此我们提出问题:分词器选择对不同语言的影响是否相同,这一问题现有文献未解答。为此,我们训练了123个语言模型,涵盖54种分词器。在主要对比中,架构、训练语料库、训练token预算和优化均固定,故模型仅在分词器上存在差异。我们发现,分词器选择对拥有更少语言模型训练数据的语言影响更大:在54种分词器中,某语言的字节每比特(BPB)的标准差随其模型训练数据占比降低而增大(31种已训练语言的斯皮尔曼相关系数ρ=-0.52,28种带词边界书写的语言ρ=-0.69)。将某语言排除在分词器训练之外,会提高我们研究的所有语言的BPB,且该惩罚对拥有更少语言模型训练数据的语言往往更大。然而,给低资源语言分配更多分词器训练数据并不会无条件帮助这些语言:等权重分配及相对于语言模型训练数据反转占比的分配方式均会提高其BPB,尤其在语言模型训练重复数据时。最后,与更好BPB相关的分词器内在属性因语言而异,进一步证明良好分词器的标准取决于语言。我们发现,量化这些属性的指标可成功用于预测下游模型的成对BPB排名,为训练语言模型前筛选分词器候选提供了实用策略。
英文摘要
Tokenizer choice affects multilingual language modeling, but vocabulary capacity is finite and vocabulary size is often constrained: improving representation for some languages often comes at the expense of others. We therefore ask whether tokenizer choice matters equally across languages, a question that the current literature leave unanswered. To this end, we train 123 language models spanning 54 tokenizers. In the main comparison, architecture, training corpus, training-token budget, and optimization are held fixed, so the models differ only in their tokenizer. We find that tokenizer choice matters more for languages with less language-model training data: across the 54 tokenizers, the standard deviation of a language's bits-per-byte (BPB) increases as its model training-data share decreases (Spearman rho = -0.52 over the 31 trained languages and -0.69 over the 28 written with word boundaries). Leaving a language out of tokenizer training raises its BPB in every language we study, and the penalty tends to be larger for languages with less language-model training data. Giving lower-resource languages a larger share of tokenizer-training data, however, does not unconditionally help those languages: both equal weighting and an allocation inverting the shares with respect to the language model training data increase their BPB, particularly when language-model training repeats data. Finally, which intrinsic tokenizer properties are associated with better BPB differs across languages, providing further evidence that what makes a good tokenizer depends on the language. We find that the metrics quantifying these properties can be successfully used to predict downstream models' pairwise BPB rankings, suggesting a practical strategy for screening tokenizer candidates before training language models.