arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

低秩摩擦实现内存高效的Transformer预训练

Low-Rank Friction for Memory-Efficient Transformer Pretraining

Rajit Rajpal, Benedict Leimkuhler

arXiv 2609.30342首次发表:更新:

发表机构

School of Mathematics; University of Edinburgh(数学学院; 爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出秩-1 iKFAD(R-iKFAD)优化器,通过低秩摩擦分解将内存开销从O(mn)降至O(m+n),在保持与iKFAD相当性能的同时几乎减半内存占用,并首次给出连续时间下无需正阻尼的收敛速率证明。

AI 中文摘要

iKFAD是一种最近提出的优化器,它用动量动力学中的自适应摩擦取代了自适应学习率,但其性能与Adam相当。其局限性在于,完整的摩擦张量$\xi\in\mathbb{R}^{m\times n}$每层携带与Adam的二阶矩缓冲区相同的$\mathcal{O}(mn)$内存开销。在此,我们将iKFAD的摩擦张量$\xi$替换为基于行和列动量统计构建的秩-1外积分解,从而得到秩-1 iKFAD(R-iKFAD)。这将每层的摩擦内存占用从$\mathcal{O}(mn)$减少到$\mathcal{O}(m+n)$,大约将iKFAD的总优化器状态减半。尽管有所减少,R-iKFAD在性能上与iKFAD保持相当:在GPT2-Nano、TinyViT、DistilBERT和GPT2-S上的实验证实,它在内存占用几乎减半的同时,匹配或超越了iKFAD,并且对超参数保持相当的鲁棒性。我们分析了两种阻尼机制下的连续时间动力学。对于线性阻尼($\gamma>0$),我们在强凸性下证明了指数收敛。对于$\gamma=0$(我们实验中的首选选项),摩擦完全由过去的动量产生,并随着动量消失而关闭,因此无法证明几何收敛。尽管如此,我们证明了收敛到极小值点,并给出了能量的匹配上下界:当正则化尺度$\epsilon_{\mathrm{stab}}$为零时,阶数为$t^{-1}$;当其为正时,阶数为$t^{-1/2}$。据我们所知,这是秩-1分解优化器在连续时间中的首个收敛速率结果,也是首个不需要正阻尼的此类结果。

英文摘要

iKFAD is a recently proposed optimiser that replaces adaptive learning rates with adaptive friction in the momentum dynamics, yet performs as well as Adam. Its limitation is that the full friction tensor $ξ\in\mathbb{R}^{m\times n}$ carries the same $\mathcal{O}(mn)$ memory overhead per layer as Adam's second-moment buffer. Here we replace iKFAD's friction tensor $ξ$ with a rank-1 outer-product factorisation built from row and column momentum statistics, resulting in Rank-1 iKFAD (R-iKFAD). This reduces the friction memory footprint from $\mathcal{O}(mn)$ to $\mathcal{O}(m+n)$ per layer, which approximately halves iKFAD's total optimiser state. Despite this reduction, R-iKFAD maintains parity in performance with iKFAD: experiments on GPT2-Nano, TinyViT, DistilBERT and GPT2-S confirm that it matches or exceeds iKFAD while nearly halving the memory footprint and remaining comparably robust to hyperparameters. We analyse the continuous-time dynamics in two damping regimes. For linear damping ($γ>0$) we prove exponential convergence under strong convexity. For $γ=0$, the preferred option in our experiments, the friction is generated entirely from past momentum and switches off as the momentum vanishes, so geometric convergence cannot be shown. We nonetheless prove convergence to the minimiser, together with matching upper and lower bounds on the energy: of order $t^{-1}$ when the regularisation scale $ε_{\mathrm{stab}}$ is zero, and of order $t^{-1/2}$ when it is positive. To our knowledge this is the first convergence rate for a rank-1 factored optimiser in continuous time, and the first such result that does not require positive damping.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑