发表机构
Peking University; Nanjing University(北京大学; 南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SOLAR提出状态驱动在线学习率调度框架,通过基础锚定残差修正和断路器机制,在LLM预训练中稳定自适应LR,提升困惑度并支持跨规模重用。
AI 中文摘要
学习率(LR)调度在大语言模型(LLM)预训练中扮演核心角色,然而当前实践仍严重依赖手工设计的启发式方法,如Warmup-Cosine-Decay和Warmup-Stable-Decay。由于这些调度是预先固定的,它们无法适应不断变化的优化动态。在“学习优化”(L2O)框架内的在线学习调度提供了一种动态替代方案,但在LLM规模下由于噪声信号、延迟反馈以及灾难性发散的风险而显得脆弱。我们提出SOLAR(状态驱动在线学习率调度器),一个用于可靠在线LR自适应的稳定框架。SOLAR以基础调度为参考,并为各个参数组学习有界、状态相关的残差修正。每个修正每一步都重新锚定到基础调度,使策略能够在不重新学习warmup-decay轮廓的情况下调整LR。轻量级状态表示和进度感知奖励指导在线学习,而断路器在罕见的不安全动作后恢复训练。在自回归语言模型预训练中,SOLAR在密集模型(从60M到1B,使用AdamW和Muon)以及两种MoE设置(最高3B)上,相比调优的静态调度和自动LR调优器,改善了最终困惑度。匹配的130M对照实验表明,添加基础锚定和动作界限将全局PPO控制器从27.09改善到23.74的最终PPL,而分组控制在相同的两个种子上达到22.87。在60M代理上训练的残差策略也可以冻结并在更大密集规模上重用,无需目标PPO更新,在四倍基础LR范围内保持有效。这些结果确立了SOLAR作为LLM预训练中实用的学习型LR控制器。
英文摘要
Learning-rate (LR) scheduling plays a central role in large language model (LLM) pretraining, yet current practice still relies heavily on hand-crafted heuristics such as Warmup-Cosine-Decay and Warmup-Stable-Decay. Because these schedules are fixed in advance, they cannot adapt to evolving optimization dynamics. Online learned scheduling within the Learning to Optimize (L2O) framework offers a dynamic alternative, but remains brittle at LLM scale due to noisy signals, delayed feedback, and the risk of catastrophic divergence. We propose SOLAR (State-driven Online Learning rAte scheduleR), a stabilized framework for reliable online LR adaptation. SOLAR uses a base schedule as a reference and learns bounded, state-dependent residual corrections for individual parameter groups. Each correction re-anchors to the base at every step, allowing the policy to adapt the LR without relearning the warmup-decay profile. A lightweight state representation and progress-aware reward guide online learning, while a Circuit-Breaker restores training after rare unsafe actions. Across autoregressive language-model pretraining, SOLAR improves final perplexity over tuned static schedules and automatic LR tuners for dense models from 60M to 1B, AdamW and Muon, and two MoE settings up to 3B. Matched 130M controls show that adding base anchoring and action bounds improves a global PPO controller from 27.09 to 23.74 final PPL, while group-wise control reaches 22.87 on the same two seeds. A residual policy trained on a 60M proxy can also be frozen and reused at larger dense scales without target PPO updates, remaining effective across a fourfold base-LR range. These results establish SOLAR as a practical learned LR controller for LLM pretraining.