无迹可循:低秩训练中的不可识别性与优化器状态
No Subspace to Track: Non-Identifiability and Optimizer State in Low-Rank Training
浏览论文内容
中文总结 AI 辅助
研究低秩训练中优化器子空间不可识别性问题,发现不同小批量计算的子空间估计差异大,平均无法恢复其余方向。提出将刷新视为坐标变化,如LDAdam表现良好,还明确了GaLore有效原因及相关检查要点在于可重现秩k*。
中文摘要 AI 辅助
内存高效的优化器(如GaLore)通过将梯度投影到每T步重新计算的秩为r的子空间来训练大语言模型,假设该子空间是一个可跟踪的缓慢漂移对象。我们表明,除了一个小的可重现核心外,不存在这样的对象。从不相交的小批量数据在同一步骤计算的前r个子空间的两个估计值,其差异与相隔T步计算的估计值一样大(在具有r = 128的Pythia - 160M模型中,最大弦距sqrt(2r)的0.73对0.74):每次刷新时的明显旋转主要由估计器噪声主导。这在三个架构类别的四个模型家族中都成立,从70M到6.9B参数,随着规模增大而增强,在视觉Transformer中则较弱。在小批量数据中,128个方向中只有约39个是可重现的,平均不能恢复其余方向:在N折平均下,梯度的谱尾以N^(-1/4)而不是纯噪声的N^(-1/2)收缩,因此没有平均预算能使子空间定义良好。相反,将每次刷新视为Adam状态的坐标变化会有所帮助。盲目携带二阶矩比最佳旋转盲估计器差约(r - k*)/2,而一阶矩能准确通过旋转传输,这是各向同性梯度下的最优线性映射以及LDAdam使用的规则。在1B参数模型上进行40k步训练(3个种子),当beta2 = 0.999时,完整的LDAdam达到18.7的困惑度,超过了经过最佳beta2修复后的未传输GaLore(19.3);将二阶矩记忆缩短到beta2 = 0.99对刷新优化器有帮助,不过对于标准GaLore,效果较小,而全秩控制会使其逆转。一个可测量的事实,即子空间不可识别性,阐明了GaLore为何有效、哪些补丁有效以及在相信低秩假设之前要检查什么:可重现秩k*。
英文摘要
Memory-efficient optimizers such as GaLore train large language models by projecting gradients onto a rank-r subspace recomputed every T steps, assuming this subspace is a slowly drifting object that can be tracked. We show that beyond a small reproducible core, there is no such object. Two estimates of the top-r subspace computed at the same step from disjoint minibatches disagree as much as estimates computed T steps apart (0.73 vs 0.74 of the maximal chordal distance sqrt(2r), at Pythia-160M with r=128): the apparent rotation at each refresh is dominated by estimator noise. This holds across four model families in three architecture classes from 70M to 6.9B parameters, strengthening with scale, and more weakly in a vision transformer. Only ~39 of 128 directions are reproducible across minibatches, and averaging cannot recover the rest: under N-fold averaging the gradient's spectral tail shrinks as N^(-1/4) rather than the N^(-1/2) of pure noise, so no averaging budget makes the subspace well defined. What helps instead follows from treating each refresh as a change of coordinates for Adam's state. Carrying the second moment blindly is provably about (r-k*)/2 worse than the best rotation-blind estimator, while the first moment transports exactly through the rotation, the optimal linear map under isotropic gradients and the rule LDAdam uses. At 1B over 40k steps (3 seeds), full LDAdam reaches 18.7 perplexity at beta2=0.999, beating untransported GaLore after its best beta2 fix (19.3); shortening the second-moment memory to beta2=0.99 helps the refreshing optimizers, though for canonical GaLore the effect is small and a full-rank control reverses it. One measurable fact, subspace non-identifiability, clarifies why GaLore works, which patches work, and what to check before trusting a low-rank assumption: the reproducible rank k*.
发表机构
- Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。