arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.21716cs.LGcs.AImath.OCstat.ML

二次模型的辩护

A Defense of the Quadratic Model

Alexandru Meterez, Pranav Ajit Nair, Depen Morwani, Cengiz Pehlevan, Sham Kakade, Alex Damian

首次发表
浏览论文内容

中文总结 AI 辅助

研究对二次模型进行压力测试,通过泰勒展开预测优化动态,利用兰索斯求积法分析海森矩阵谱和局部稳定性,发现其在语言模型中能准确反映优化情况,或可作为预训练优化动态的理论易处理代理。

中文摘要 AI 辅助

由于神经网络损失景观的复杂性,优化理论不得不依赖理想化模型,且在模型的理论易处理性与其对真实优化动态的描述准确性之间通常存在权衡。在这项工作中,我们对最简单的优化模型——二次模型进行压力测试,表明它在具有1.5亿参数和30亿训练令牌的语言模型设置中具有惊人的预测能力。具体而言,通过在训练过程中的中间检查点对模型和损失函数进行泰勒展开,可以准确预测长达10%训练时长的窗口内的优化动态。在确立了这种一致性之后,我们通过两个视角分析这些局部二次优化问题的结构:海森矩阵谱和局部稳定性。使用具有极深探针的兰索斯求积法,我们能够深入估计海森矩阵谱的尾部,并且在特征值和特征向量中发现了大量结构,这取决于批量大小、预处理器和训练时间。我们还在中间检查点对局部线性稳定性进行了实证测试,并将其与理论预测进行比较,以证明语言模型中的优化通常发生在随机稳定性边缘,其性质也由批量大小决定。我们的结果表明,二次模型可能是预训练优化动态的理论上易处理的代理。

英文摘要

Due to the complexity of neural network loss landscapes, optimization theory is forced to rely on idealized models, and there is generally a tradeoff between how theoretically tractable the model is, and how accurately it describes the true optimization dynamics. In this work, we stress test the simplest possible model of optimization -- the quadratic model -- and show that it can be surprisingly predictive in an LLM setting with 150M parameters and 3B training tokens. Specifically, we show that Taylor expanding the model and the loss function at intermediate checkpoints through training can accurately predict the optimization dynamics over windows that can last up to 10\% of training. Having established this agreement, we then turn to analyzing the structure of these local quadratic optimization problems through two lenses: the Hessian spectrum and local stability. Using Lanczos quadrature with extremely deep probes, we are able to estimate the Hessian spectrum deep into the tail, and we find a surprising amount of structure in both the eigenvalues and eigenvectors, which depends on the batch size, preconditioner, and training time. We also empirically test local linear stability at intermediate checkpoints and compare it to theoretical predictions to demonstrate that optimization in LLMs typically occurs at a stochastic edge of stability, whose nature is also determined by batch size. Our results indicate the quadratic model may be a theoretically tractable proxy for pretraining optimization dynamics.

发表机构

  • Kempner Institute at Harvard University(哈佛大学坎普纳研究所)
  • MIT(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

↑