步进法则能否迁移到小规模语言模型?59M 参数以下的实证重新校准
Does Step Law Transfer to Small-Scale Language Models? An Empirical Recalibration Below 59M Parameters
浏览论文内容
中文总结 AI 辅助
本文实证检验步进法则在小于59M参数的小规模语言模型上的迁移性,发现幂律形式仍成立但系数需重新校准,直接迁移会高估最优学习率约4倍。
中文摘要 AI 辅助
步进法则为预训练语言模型时的最优峰值学习率 eta* 和批大小 B* 提供了幂律公式。该法则在 59M 至 1B 参数的模型上进行了校准;其作者从未在 N < 59M 的小模型范围内进行过实证测试。这一范围对于单 GPU 训练、可解释性研究、教育实验以及因内存或成本原因无法使用更大模型的场景具有重要意义。我们测试了步进法则是否适用于小语言模型。我们考虑三种结果:H1,原始系数直接适用;H2,幂律形式成立但系数不同;H3,幂律无法描述该范围内的最优值。所有实验均使用单一的 nanoGPT/TinyStories 流程,包含 2048 个 token 的 BPE 词表、AdamW 优化器和 warmup-cosine 学习率调度。每个 (N, D) 单元的最优值通过在对数-对数坐标下对平滑训练损失进行局部二次近似从损失曲面 L(eta, B) 中提取。最终数据集包含 29 个唯一的 (N, D) 单元和 935 次分析就绪的运行。主要重新拟合使用了 25 个单元(815 次运行),工作范围为 4 <= D/N <= 600。在合并数据上,我们接受 H2:函数形式得以保留,但系数与原始不同。我们得到 eta*(N, D) = 0.0985 N^(-0.508) D^(0.238)(R^2 = 0.834)和 B*(D) = 3.6 x 10^(-4) D^(0.931)(R^2 = 0.950)。步进法则关于 B* 独立于 N 的结构性主张被重现(p = 0.87),但 B* 随 D 的增长几乎是原始工作的两倍陡峭。步进法则的直接迁移系统地高估了最优学习率:中位数比值 eta_SL / eta* 约为 4.0 倍,范围为 2.4 倍至 6.6 倍。
英文摘要
Step Law gives power-law formulas for the optimal peak learning rate eta* and batch size B* when pre-training language models. It was calibrated on models between 59M and 1B parameters; the small-model regime N < 59M was never tested empirically by its authors. This regime matters for single-GPU training, interpretability research, educational experiments, and settings where larger models are infeasible on memory or cost grounds. We test whether Step Law transfers to small language models. We consider three outcomes: H1, the original coefficients work directly; H2, the power-law form holds but with different coefficients; and H3, a power law does not describe the optima in this regime. All experiments use a single nanoGPT/TinyStories pipeline with a 2048-token BPE vocabulary, AdamW, and a warmup-cosine schedule. The optimum for each (N, D) cell is extracted from the loss surface L(eta, B) via a local quadratic approximation in log-log coordinates over the smoothed training loss. The final dataset contains 29 unique (N, D) cells and 935 analysis-ready runs. The main refit uses 25 cells (815 runs) in the working range 4 <= D/N <= 600. On the pooled data we accept H2: the functional form is preserved, but the coefficients differ from the original. We obtain eta*(N, D) = 0.0985 N^(-0.508) D^(0.238) (R^2 = 0.834) and B*(D) = 3.6 x 10^(-4) D^(0.931) (R^2 = 0.950). Step Law's structural claim that B* is independent of N is reproduced (p = 0.87), but the growth of B* with D is nearly twice as steep as in the original work. Direct transfer of Step Law systematically overestimates the optimal learning rate: the median ratio eta_SL / eta* is approximately 4.0x, with a range of 2.4x to 6.6x.
发表机构
- HSE University(高等经济大学)
- Faculty of Computer Science(计算机科学学院)
机构由 AI 辅助整理,请以论文原文为准。