发表机构
Kyutai; Sorbonne Université; CNRS; Institute of Intelligent Systems and Robotics; Univ. Grenoble Alpes; Grenoble INP; LIG(Kyutai; 索邦大学; 法国国家科学研究中心; 智能系统与机器人研究所; 格勒诺布尔阿尔卑斯大学; 格勒诺布尔国立理工学院; 格勒诺布尔信息学实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多语言大模型共享词汇表导致压缩不均、内存占用高和推理慢的问题,提出模块化分词器框架,通过可提取子分词器及采样预训练策略,实现高效训练与推理,不牺牲性能。
AI 中文摘要
多语言大语言模型(LLMs)传统上依赖所有支持语言共享的单一词汇表,这可能导致不同语言间的压缩不均。此外,它们庞大的嵌入矩阵和输出矩阵增加了内存使用并减慢了推理速度,尤其对于小规模模型。这也是一种浪费,因为模型通常仅用于部分语言子集。为解决这些问题,我们引入了一个用于多语言模型训练的模块化框架。首先,我们提出了学习大型模块化BPE和Unigram分词器的方法,这些分词器能够提取适用于任何语言子集的子分词器。这些子分词器实现了与单语分词器相当的压缩率,并改善了跨语言的公平性。其次,我们设计了一种预训练策略,该策略对子分词器进行采样以形成批次,将预测限制在相关的词汇子集内,从而在词汇量大的情况下实现高效训练。这支持使用任意语言特定词汇组合进行高效推理。因此,它在不牺牲性能的情况下减少了内存使用并加速了模型推理。
英文摘要
Multilingual Large Language Models (LLMs) traditionally rely on a single vocabulary shared by all supported languages, which can lead to uneven compression across them. Moreover, their large embedding and output matrices increase memory usage and slow inference, notably for small-scale models. It is also wasteful as models are often used for only a subset of languages. To address these issues, we introduce a modular framework for multilingual model training. First, we propose methods to learn large modular BPE and Unigram tokenizers that enable extraction of subtokenizers tailored to any language subset. These subtokenizers achieve compression on par with monolingual tokenizers and improve cross-lingual fairness. Second, we design a pretraining strategy that samples subtokenizers to form batches, restricting predictions to the relevant vocabulary subset and allowing efficient training despite a large vocabulary. This supports efficient inference with any combination of language-specific vocabularies. Therefore, it reduces memory usage and speeds up inference in models without sacrificing performance.