发表机构
University of Wisconsin–Madison; Shanghai Jiao Tong University(威斯康星大学麦迪逊分校; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过理论分析和nanoGPT实验,揭示了学习率与批次大小联合调度对LLM预训练幂律学习曲线的影响,提出3+3(+2)标度律机制并验证其预测能力。
AI 中文摘要
幂律学习曲线常被视为模型及其数据的固定属性,尽管学习率和批次大小调度可以改变观测到的损失。我们在具有线性随机特征的噪声在线SGD中研究这种依赖性。以表示为条件,一个精确的Volterra方程分离出两个响应分量:一个传播未解析目标误差的强迫项和一个传播随机误差注入的记忆核。我们证明,任一分量遵循幂律当且仅当其累积加权谱质量具有相应的低谱标度;单个特征值和目标系数无需服从坐标级幂律。在联合调度下,内在时间$T_t=\sum_{s<t}\eta_s$控制优化进度,而$r_t=B_t/\eta_t$控制噪声注入。它们的相互作用产生尖锐条件,在这些条件下,调度保留、改变或破坏干净的幂律,同时存在噪声降低的记忆上限。幂律随机特征模型在$3+3(+2)$传播机制中实现该机制,具有相位相关的计算速率。受控nanoGPT实验表明:(1)具有匹配$B/\eta$路径的学习率和批次大小调度在内在时间上几乎等价,(2)一个强迫-记忆替代模型能准确预测跨调度的损失,(3)其在真实世界数据集上的拟合指数识别出$3+3(+2)$图中LLM所处的机制。
英文摘要
Power-law learning curves are often treated as fixed properties of a model and its data, although learning-rate and batch-size schedules can change the observed loss. We study this dependence in noisy online SGD with linear random features. Conditional on the representation, an exact Volterra equation separates two response components: a forcing term that propagates unresolved target error and a memory kernel that propagates stochastic-error injections. We prove that either component follows a power law if and only if its cumulative weighted spectral mass has the corresponding low-spectrum scaling; individual eigenvalues and target coefficients need not obey coordinatewise power laws. Under a joint schedule, intrinsic time $T_t=\sum_{s<t}η_s$ controls optimization progress, while $r_t=B_t/η_t$ controls noise injection. Their interaction yields sharp conditions under which a schedule preserves, changes, or destroys the clean power law, together with a memory ceiling on noise reduction. The power-law random-feature model realizes this mechanism in $3+3(+2)$ propagation regimes with phase-dependent compute rates. Controlled nanoGPT experiments show that (1) learning-rate and batch-size schedules with matched $B/η$ paths are nearly equivalent in intrinsic time, (2) a forcing-memory surrogate accurately predicts loss across schedules, and (3) its fitted exponents across real-world datasets identify the regime of LLMs in $3+3(+2)$ map.