发表机构
Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究预训练语言模型中反复出现的谱模式能否用于初始化,通过分析多个检查点构建初始化方案并与其他方法比较,发现预训练权重复用有竞争力,但仅粗略谱匹配非可靠策略,预训练谱是有用诊断,有效复用需保留更多信息。
AI 中文摘要
预训练语言模型常常呈现出结构化的权重谱,这表明训练可能会反复产生相似的分层和逐组件组织。我们探讨这些反复出现的谱模式是否可作为GPT-2风格语言模型预训练的初始化信号。首先,分析了11个在大小、语言、分词器和训练语料库方面各异的预训练GPT-2风格检查点,测量各层和Transformer子组件的弗罗贝尼乌斯范数和有效秩熵。检查点呈现出共享的深度趋势。接着构建模仿预训练模型逐组件大小和谱轮廓的初始化方案,并与多种权重初始化方法比较。这些初始化器改变了模型的结构谱模式,但评估结果未显示出相应的性能优势。预训练权重复用仍具竞争力,仅粗略的谱匹配不是可靠的优化策略。结果表明预训练谱对训练模型结构是有用的诊断,但有效的复用可能需要保留比分层大小和奇异值形状更丰富的信息。
英文摘要
Pretrained language models often exhibit structured weight spectra, suggesting that training may repeatedly produce similar layerwise and component-wise organization. We ask whether these recurring spectral patterns can be reused as an initialization signal for GPT-2-style language-model pretraining. First, we analyze eleven pretrained GPT-2-style checkpoints that vary in size, language, tokenizer, and training corpus, measuring Frobenius norm and effective-rank entropy across layers and Transformer subcomponents. The checkpoints show shared depth trends, especially increasing scale and stronger spectral concentration in residual-writing matrices. We then construct initialization schemes that imitate the component-wise magnitudes and spectral profiles of pretrained models, and compare them with several weight initialization methods. These initializers visibly change the model's structural spectral patterns, but the evaluation results do not show a corresponding performance advantage. Pretrained-weight reuse remains competitive, while coarse spectral matching alone is not a reliable optimization strategy. Our results suggest that pretrained spectra are useful diagnostics of trained model structure, but that effective reuse likely requires preserving richer information than component-wise scale and singular-value shape.