arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27680cs.LGcs.CL

低秩适配的紧样本复杂度:匹配界与秩选择

Tight Sample Complexity for Low-Rank Adaptation: Matching Bounds and Rank Selection

Arunan J

首次发表
浏览论文内容

中文总结 AI 辅助

本文解决LoRA微调的样本复杂度问题,建立匹配的上下界,得出秩选择二分性,在合成与真实基准上验证理论预测,揭示过参数化惩罚的来源。

中文摘要 AI 辅助

低秩适配(Low-Rank Adaptation, LoRA)已成为对大型预训练模型进行微调的标准机制,但其统计特性仍仅被部分理解。现有泛化结果提供了形如Õ(√(rd/n))或Õ(rd/n)的上界,但缺失匹配的下界,且如何选择LoRA秩r的问题尚无正式答案。本文填补了这两个空白:采用局部拉德马赫论证建立了经验风险最小化器在秩为r的LoRA上的超额风险的上界Õ(rd/n),当目标适配的秩最多为r时成立;随后通过对R^{d×d}的秩为r子空间进行Fano型打包,证明了匹配的极小极大下界Ω(rd/n),该下界适用于输出位于秩为r的LoRA类中的任何估计器。结合两者得到秩选择二分性:对于约束经验风险最小化器,最优秩等于固有秩r*,过度选择秩会严格损害性能;对于核范数后截断型自适应估计器,过度选择秩无害,且速率饱和为Θ̃(r*d/n),与r无关。综上,这三个结果在良好指定的局部二次框架内刻画了LoRA微调的统计复杂度,并将经验观察到的过参数化惩罚识别为未正则化经验风险最小化的特性,而非LoRA类本身的特性。该理论的预测在合成轨迹回归基准及真实LoRA微调任务上得到验证,涉及DistilBERT和RoBERTa在SST-2和MRPC上的三个(模型,任务)配置,所有配置均呈现预测的验证损失U型曲线,其中两个配置在大秩下表现出统计显著的损失膨胀(配对置换p=0.016)。

英文摘要

Low-Rank Adaptation (LoRA) has become the standard mechanism for fine-tuning large pretrained models, yet its statistical properties remain only partially understood. Existing generalization results provide upper bounds of the form O~(sqrt(rd/n)) or O~(rd/n), but a matching lower bound is missing, and the question of how to choose the LoRA rank r has no formal answer. Both gaps are closed here. A local Rademacher argument establishes an upper bound of O~(rd/n) on the excess risk of the empirical risk minimizer over rank-r LoRA, whenever the target adaptation has rank at most r. A matching minimax lower bound of Omega(rd/n) is then proved via a Fano-type packing of the rank-r subspace of R^{d x d}; the bound applies to any estimator whose output lies in the rank-r LoRA class. Combining the two yields a rank-selection dichotomy. For the constrained empirical risk minimizer, the optimal rank equals the intrinsic rank r*, and over-ranking strictly hurts. For adaptive estimators of the nuclear-norm-then-truncate type, over-ranking is harmless and the rate saturates at Theta~(r* d / n) regardless of r. Taken together, the three results characterize the statistical complexity of LoRA fine-tuning within the well-specified locally quadratic regime, and identify the empirically observed over-parameterization penalty as a property of unregularized empirical risk minimization rather than of the LoRA class itself. Predictions of the theory are verified on a synthetic trace-regression benchmark and on real LoRA fine-tuning across three (model, task) configurations covering DistilBERT and RoBERTa on SST-2 and MRPC. All configurations exhibit the predicted U-shape in validation loss, with two showing statistically significant loss inflation at large ranks (paired permutation p = 0.016).

补充信息

↑