发表机构
The Chinese University of Hong Kong; Westlake University(香港中文大学; 西湖大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究了随机低秩适应的收敛性,提出LoRA-NSGDM和LoRA-STORM算法,分别在$\mathcal{O}(\epsilon^{-8})$和$\mathcal{O}(\epsilon^{-6})$的随机oracle复杂度下找到$\epsilon$-stationary点。
AI 中文摘要
低秩适应(LoRA)在两个适配器 $B \in \mathbb{R}^{m \times r}$ 和 $A \in \mathbb{R}^{r \times n}$ 上优化 $J(B,A)=\mathcal L(W_\mathrm{base}+sBA)$,其中 $W_\mathrm{base} \in \mathbb{R}^{m \times n}$ 是冻结的预训练权重矩阵。先前的分析表明,在确定性设置下,LoRA-GD需要 $\exp\{\mathcal{O}(\epsilon^{-2})\}$ 个oracle调用来找到一个 $\epsilon$-stationary点,使得 $\|\nabla J(B,A)\|\leq \epsilon$。我们改进了分析,证明对于相同的首阶标准,$\mathcal{O}(\epsilon^{-4})$ 次完整梯度评估就足够。我们进一步研究了在无偏梯度估计和有限方差下的随机LoRA。我们提出了LoRA-NSGDM,它在 $\mathcal{O}(\epsilon^{-8})$ 次随机oracle复杂度下找到 $\epsilon$-stationary点。在额外的均方光滑性条件下,我们使用方差减少策略并提出LoRA-STORM,将随机oracle复杂度提升到 $\mathcal{O}(\epsilon^{-6})$。
英文摘要
Low-rank adaptation (LoRA) optimizes $J(B,A)=\mathcal L(W_\mathrm{base}+sBA)$ over two adapters $B \in \mathbb{R}^{m \times r}$ and $A \in \mathbb{R}^{r \times n}$ that form a low-rank update to a frozen pretrained weight matrix $W_\mathrm{base} \in \mathbb{R}^{m \times n}$. The prior analysis shows LoRA-GD takes $\exp\{\mathcal{O}(ε^{-2})\}$ oracle calls to find an $ε$-stationary point such that $\|\nabla J(B,A)\|\leq ε$ in the deterministic setting. We sharpen the analysis and show that $\mathcal{O}(ε^{-4})$ full-gradient evaluations suffice for the same first-order criterion. We further study stochastic LoRA under unbiased gradient estimates and finite variance. We propose LoRA-NSGDM, which finds an $ε$-stationary point with $\mathcal{O}(ε^{-8})$ stochastic oracle complexity. Under the additional mean-square smoothness condition, we use variance reduction strategy and propose LoRA-STORM, which improves the stochastic oracle complexity to $\mathcal{O}(ε^{-6})$.