arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越初始化损失:大语言模型词汇扩展的词元嵌入初始化策略的系统研究

Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension

Raviraj Joshi, Utkarsh Vaidya, Sanjay Singh Chauhan, Niranjan Wartikar

arXiv 2608.03494首次发表:更新:

AI 中文总结

该研究针对LLM词汇扩展,对比20余种初始化策略,发现子词组合方法更优,最优配置使CPT步骤减超6倍,轻量CPT可可靠选择最优初始化策略。

AI 中文摘要

词汇扩展是使预训练大语言模型(LLMs)适应新语言的有效方式,但新增词元嵌入的初始化会强烈影响持续预训练(CPT)的效率。我们对Nemotron-3-Nano-30B-A3B的印地语词汇扩展开展了超过20种初始化策略的系统研究,对比范围涵盖词汇平均基线、外部及学习型初始化方法(包括FOCUS、top-k语义检索、残差MLP映射)、子词组合、范数校准以及输入输出不对称性。研究发现,子词组合方法的性能优于词汇平均方法和外部/学习型初始化方法;在子词组合中,非对称变体实现了目前观测到的最低早期验证损失,且呈现出输入与输出嵌入初始化的不同偏好。最优配置将输入嵌入矩阵初始化为均匀子词平均并采用印地语特定范数校准,将输出语言建模头初始化为字符长度加权的子词平均。相较于标准Mean-all基线,该完整初始化管线达到可比验证损失的同时,CPT步骤减少超过6倍,且仅在500步后就超过了基线在3500步时的MILU-Hindi准确率。最后,我们表明初始化损失和初始化每字节比特数(Init BPB)是下游收敛的不可靠预测指标,而仅需50步的轻量CPT为选择最优初始化策略提供了高性价比且可靠的信号。

英文摘要

Vocabulary extension is an efficient way to adapt pretrained large language models (LLMs) to new languages, but the initialization of newly added token embeddings can strongly affect continued pre-training (CPT) efficiency. We present a systematic study of more than 20 initialization strategies for Hindi vocabulary extension in Nemotron-3-Nano-30B-A3B. Our comparison spans vocabulary-averaging baselines; external and learned initialization methods, including FOCUS, top-k semantic retrieval, and residual MLP mappings; subword composition; norm calibration; and input-output asymmetry. We find that subword composition methods outperform both vocabulary averaging and external/learned initialization approaches. Within subword composition, asymmetric variants achieve the lowest observed early validation loss and reveal distinct preferences for input and output embedding initialization. The best observed configuration initializes the input embedding matrix with uniform subword averaging and Hindi-specific norm calibration, and the output language modeling head with character-length-weighted subword averaging. Relative to the standard Mean-all baseline, this full initialization pipeline reaches comparable validation loss with over a 6x reduction in CPT steps and exceeds the baseline's 3,500-step MILU-Hindi accuracy after only 500 steps. Finally, we show that initialization loss and initialization bits-per-byte (Init BPB) are unreliable predictors of downstream convergence, whereas lightweight CPT, as few as 50 steps, provides a cost-effective and reliable signal for selecting the best initialization strategy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑