arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Stiefel-AdamW:用于线性分解块的几何感知AdamW

Stiefel-AdamW: Geometry-Aware AdamW for Linear Factorization Blocks

Emanuele Zangrando, Marco Sutti, Francesco Tudisco

arXiv 2609.21039首次发表:更新:

发表机构

Gran Sasso Science Institute; University of Edinburgh(格兰萨索科学研究所; 爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对线性分解块中分解不唯一导致的训练不稳定问题,提出Stiefel-AdamW优化器,通过将一因子约束于Stiefel流形以消除规范对称性,在保持AdamW效率的同时提升稳定性,并在多种模型上验证了其有效性。

AI 中文摘要

现代深度学习中的一个普遍结构模式是线性分解块:一种形式为$W = BA$的子模块,其中两个参数矩阵直接相乘,中间没有非线性干预。此类块出现在LoRA适配器、低秩压缩层、自注意力的查询-键乘积中,并共享一个共同的病理:分解不唯一,这可能破坏训练稳定性并限制可用的学习率。尽管如此,分解块通常使用忽略底层几何的标准欧几里得方法进行优化。我们引入了Stiefel-AdamW,一种近乎即插即用的AdamW替代品,用于任何出现此类块的地方。通过将一个因子约束在Stiefel流形上,同时保持另一个因子为欧几里得,Stiefel-AdamW将完整的$\mathrm{GL}(\mathbb{R}^r)$规范对称性松弛为紧致正交对称性,排除了因子爆炸,同时保留了赋予AdamW实际优势的逐坐标对角预条件。矩估计在环境欧几里得空间中进行,几何仅通过切空间投影和流形回缩进入。与AdamW相比,实现开销极小,我们证明了所得到的优化器继承了黎曼方法的稳定性优势和标准收敛保证。我们在GPT2、ViT和Mistral 7B的LoRA风格微调以及GPT2在OpenWebText上的完整预训练上验证了Stiefel-AdamW,显示出在几乎不增加AdamW额外成本的情况下,相对于强基线的持续改进。

英文摘要

A pervasive structural pattern in modern deep learning is the linear factorization block: a submodule of the form $W = BA$ in which two parameter matrices are multiplied directly, with no intervening nonlinearity. Such blocks appear in LoRA adapters, low-rank compressed layers, query-key products of self-attention, and share a common pathology: the factorization is non-unique, which can destabilize training and limit usable learning rates. Despite this, factorization blocks are typically optimized with standard Euclidean methods that ignore the underlying geometry. We introduce Stiefel-AdamW, a near drop-in replacement for AdamW for use wherever such blocks appear. By constraining one factor on the Stiefel manifold while leaving the other Euclidean, Stiefel-AdamW relaxes the full $\mathrm{GL}(\mathbb{R}^r)$ gauge symmetry to a compact orthogonal symmetry, ruling out factor blow-up while retaining the coordinate-wise diagonal preconditioning that gives AdamW its practical strength. Moment estimation is performed in the ambient Euclidean space, with geometry entering only through a tangent-space projection and a manifold retraction. The implementation overhead over AdamW is minimal, and we show that the resulting optimizer inherits both the stability benefits of Riemannian methods and standard convergence guarantees. We validate Stiefel-AdamW on LoRA-style fine-tuning of GPT2, ViT, and Mistral 7B and on full pretraining of GPT2 on OpenWebText, showing consistent improvements over strong baselines at essentially no additional cost over AdamW.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑