稳定动态低秩训练
Stabilizing the Dynamic Low-Rank Training
浏览论文内容
中文总结 AI 辅助
针对动态低秩训练在高压缩下失效的问题,提出SDLRT方法,通过补偿缓冲区重注入被忽略的奇异方向并引入截断容差负反馈,在SuperGLUE上以2.8%参数开销超越LoRA。
中文摘要 AI 辅助
直接在低秩参数化中训练神经网络是一种在训练和推理过程中同时减少内存、计算和存储的有吸引力的途径。动态低秩训练(DLRT)通过梯度流的Galerkin投影将权重限制在秩为$r$的流形上,因其无需专门初始化或后分解即可在训练过程中识别高效子网络而特别有吸引力。然而,DLRT在高压缩率下无法找到可训练的网络。在本文中,我们推导了最佳秩-$r$近似的梯度流,并指出DLRT的偏移源于一个曲率耦合项,该在激进压缩下数值较大且不可忽略。基于此分析,我们提出了一种稳定的动态低秩训练方法,命名为SDLRT,它维护一个轻量级补偿缓冲区,用于重新注入被忽略的主要奇异方向。此外,我们引入对截断容差的负反馈以稳定每层的秩。实验上,SDLRT在DLRT失效的情况下可靠地找到可训练的子网络,并且作为DeBERTa-v3上的PEFT适配器,在SuperGLUE上以仅$2.8\%$的参数开销(相对于LoRA)取得了最佳平均分数。
英文摘要
Training neural networks directly in a low-rank parameterization is an appealing route to reducing memory, compute, and storage simultaneously during both training and inference. Dynamic low-rank training (DLRT), which confines weights to a rank-$r$ manifold via the Galerkin projection of the gradient flow, is particularly attractive because it identifies efficient subnetworks on the fly without specialized initialization or post-factorization. However, DLRT fails to find trainable networks under high compression. In this paper, we derive the gradient flow of the best rank-$r$ approximation and point out that the offset of DLRT comes from a curvature-coupling term which is large and thus non-negligible under aggressive compression. Guided by this analysis, we propose a stable dynamic low-rank training method, named SDLRT, which maintains a lightweight compensation buffer that reinjects the top neglected singular directions. Additionally, we introduce a negative feedback on the truncation tolerance to stabilize each layer's rank. Experimentally, SDLRT reliably finds trainable subnetworks where DLRT collapses and as a PEFT adapter on DeBERTa-v3, it achieves the best average score on SuperGLUE at only $2.8\%$ parameter overhead over LoRA.