发表机构
University of Wisconsin – Madison; Baidu Inc.; University of Arizona; University of Toronto; Wilfrid Laurier University; York University(威斯康星大学麦迪逊分校; 百度公司; 亚利桑那大学; 多伦多大学; 威尔弗里德·劳里埃大学; 约克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出二次精确合并(QAM)方法,通过检查点索引矩条件实现与序列参考二阶一致的合并,证明固定长度GD历史无法达到o(h^3)端点误差,而QAM达到最优O(h^3)界,并在真实模型上验证其长窗口优势。
AI 中文摘要
保存的检查点记录了训练轨迹中的状态,但通常不能确定在不同调度下将访问的状态处的更新。我们研究了这些检查点在给定更新强度下重建序列参考端点的准确性。在常见的局部转移模型下,两个检查点索引矩条件刻画了所有与该参考在二阶上一致的凸合并。然后我们证明了一个信息极限:对于非退化配置,仅使用步长为 $h$ 的固定长度梯度下降(GD)历史记录的任何算法,都无法在固定类光滑强凸损失上实现均匀的 $o(h^3)$ 端点误差。该下界源于两个具有相同GD检查点历史但序列参考端点相差 $\Omega(h^3)$ 的损失。\textbf{二次精确合并}(QAM)实现了匹配的均匀 $O(h^3)$ 端点误差界。其显式系数还定义了唯一的依赖于配置的合并,该合并精确匹配所有固定二次目标上的序列GD参考。在两条公开的Adam检查点轨迹(SmolLM3-3B和OpenEuroLLM-Prelude-9B)上,每个模型三个窗口和三个配置,以及15个任务中,QAM在短窗口上表现混合,在较长窗口上相对于\textbf{热启动稳定与合并}(WSM)显示出更广泛的优势。匹配矩的GSM8K诊断进一步表明,仅局部一致性并不能完全决定下游分数。这些结果刻画了保存历史的重建极限,提供了一种达到最优速率的系数规则,并评估了其实用性。
英文摘要
Saved checkpoints record states along a training trajectory, but generally do not determine the updates at states that would be visited under a different schedule. We study how accurately these checkpoints can reconstruct the endpoint of a sequential reference with prescribed update strengths. Under a common local transition model, two checkpoint-index moment conditions characterize all convex merges that agree with this reference through second order. We then prove an information limit that for nondegenerate profiles, no algorithm using only a fixed-length gradient-descent (GD) history with step size $h$ can achieve $o(h^3)$ endpoint error uniformly over a fixed class of smooth, strongly convex losses. The lower bound follows from two losses with identical GD checkpoint histories but sequential reference endpoints separated by $Ω(h^3)$. \textbf{Quadratic-Accurate Merging} (QAM) achieves a matching uniform $O(h^3)$ endpoint error bound. Its explicit coefficients also define the unique profile-dependent merge that exactly matches the sequential GD reference across all fixed quadratic objectives. Across two public Adam checkpoint trajectories (SmolLM3-3B and OpenEuroLLM-Prelude-9B), three windows and three profiles per model, and 15 tasks, QAM shows mixed results for short windows and broader advantages over \textbf{Warmup-Stable and Merge} (WSM) for longer windows. Matched-moment GSM8K diagnostics further show that local consistency alone does not fully determine downstream scores. These results characterize the reconstruction limits of saved histories, provide a coefficient rule that attains the optimal rate, and assess its practical utility.