arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

预训练语言模型的就地分词器扩展

In-Place Tokenizer Expansion for Pre-trained LLMs

Jimmy T. H. Smith, Tarek Dakhran, Alberto Cabrera, Simon S. Lee, Paul Pak, Aditya Tadimeti, Tim Seyde, Maxime Labonne, Alexander Amini, Mathias Lechner

arXiv 2607.15232首次发表:更新:

AI 中文总结

研究针对预训练语言模型分词器问题,提出就地分词器扩展方法。在多语言语料库上延续字节对编码合并,经两阶段适应恢复质量。应用于LFM2 - 8B - A1B生成LFM2.5 - 8B - A1B,提升多种语言解码速度,还发布模型权重等并报告负面发现。

AI 中文摘要

预训练开始时的固定分词器按预训练语料库比例分配词汇,反映当时的部署优先级。当优先级变化时,后添加的语言每个单词会被拆分成更多词元,增加延迟、计算和能耗。云模型能容纳广泛词汇,而紧凑型模型中嵌入和语言模型头矩阵占每个词元解码带宽的很大部分,所以设备端模型词汇量小且接受固定语言集外的碎片化。我们提出分词器扩展,这是一种在模型生产者控制设计时就地升级预训练模型分词器的方法。我们在多语言语料库上延续现有分词器的字节对编码合并,大部分源词元作为单个词元不变,每个新词元可精确分解为源词元。复制不变的嵌入行,将新行初始化为源子词元嵌入的均值。通过两阶段适应,即仅嵌入训练然后全模型继续预训练,恢复源检查点质量。我们将此方法应用于LFM2 - 8B - A1B(一个8B参数的专家混合模型)的继续预训练检查点,以帮助生成具有128K分词器的LFM2.5 - 8B - A1B。扩展后的分词器对印地语和越南语的编码词元比源分词器少约2.4倍和2.6倍(泰语高达4.0倍)。结合这些减少量和更大词汇量的每个词元成本测量,我们估计在我们的参考设备上这些语言的每个字符解码速度提高2.2 - 3.7倍。我们发布了模型权重和扩展后的分词器,并报告了形成该方法的负面发现。

英文摘要

A tokenizer fixed at the start of pre-training allocates vocabulary in proportion to the pre-training corpus, reflecting the deployment priorities at that time. When those priorities shift, languages added later are split into many more tokens per word, which can raise latency, compute, and energy consumption for users of those languages. Cloud models can afford a broad vocabulary because the embedding and LM-head matrices are a small fraction of their parameters. On a compact model those matrices are a material share of per-token decode bandwidth, so on-device models ship small vocabularies and accept fragmentation outside a fixed language set. We present tokenizer expansion, an in-place recipe for upgrading a pre-trained model's tokenizer when the model producer controls its design. We continue the existing tokenizer's BPE merges on a multilingual corpus, so most source tokens carry over unchanged as single tokens and every new token has an exact decomposition into source tokens. We copy the carried-over embedding rows unchanged and initialize new rows as the mean of their source sub-token embeddings. A two-stage adaptation, embedding-only training then full-model continued pre-training, recovers source-checkpoint quality. We apply the recipe to a continued pre-trained checkpoint of LFM2-8B-A1B, an 8B-parameter Mixture-of-Experts model, to help produce LFM2.5-8B-A1B with a 128K tokenizer. The expanded tokenizer encodes Hindi and Vietnamese in roughly $2.4\times$ and $2.6\times$ fewer tokens than the source (up to $4.0\times$ on Thai). Combining these reductions with the measured per-token cost of the larger vocabulary, we estimate a $2.2$-$3.7\times$ per-character decode speedup for these languages across our reference devices. We release the model weights and the expanded tokenizer, and report the negative findings that shaped the recipe.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑