秩买到了什么?低秩适应的谱分析与分布分析
What Does the Rank Buy? A Spectral and Distributional Analysis of Low-Rank Adaptation
浏览论文内容
中文总结 AI 辅助
本文通过谱与分布分析证明,在硬性范数预算下LoRA的秩不控制容量,而决定可达更新与抵消预训练方向的代价。
中文摘要 AI 辅助
LoRA中的秩 $r$ 被广泛视为一种容量控制:较小的秩被认为会产生更简单的模型,从而具有更好的泛化性能。我们证明,在硬性的逐因子范数预算下——这是实践中使用的权重衰减和范数控制的理想化形式——这种直觉会失效。原因在于结构性的:在这样的预算下,LoRA所能达到的更新恰好是核范数球内秩至多为 $r$ 的矩阵,而我们分析的每一个复杂度泛函和位移泛函都在这个集合上被秩一更新最大化——因此秩上限从未起到约束作用。其后果直接显现。我们所研究的线性读出模型类对于每个 $r \ge 1$ 都是相同的,其Rademacher复杂度不依赖于 $r$,并且适应过程能使源分布移动的距离满足一个与秩无关的上界,我们证明该上界是紧的。如果秩不能控制容量,那它在哪里起作用?我们确定了两个位置。在统计上,将逐因子预算替换为对乘积的联合预算,可以恢复一个依赖于数据、对秩敏感的复杂度界——尽管这种增益只出现在特征分布充分分散的情况下,而最坏情况仍然与秩无关。在谱上,秩决定了适应的代价:抵消预训练权重的主要奇异方向既需要足够的秩,也需要足够的预算。我们给出了实现期望的源-目标对齐所需的最小秩的上下界,在双侧谱衰减下,这两个界是匹配的。综合来看,这些结果将秩重新定义为决定哪些更新可达以及抵消的代价——而不是模型拥有多少容量。
英文摘要
The rank $r$ in LoRA is widely treated as a capacity control: a smaller rank is assumed to yield a simpler model that generalizes better. We show that, under hard per-factor norm budgets---the idealization of the weight decay and norm control used in practice---this intuition breaks down. The reason is structural: under such budgets, the updates LoRA can reach are exactly the matrices of rank at most $r$ inside a nuclear-norm ball, and every complexity and displacement functional we analyze is maximized over this set by a rank-one update---so the rank cap never binds. The consequences follow directly. The linear-readout model class we study is identical for every $r \ge 1$, its Rademacher complexity carries no dependence on $r$, and the distance the adaptation can move the source distribution obeys a rank-independent upper bound that we show is sharp. If rank does not control capacity, where does it act? We identify two places. Statistically, replacing the per-factor budgets with a joint budget on the product restores a data-dependent, rank-sensitive complexity bound---though the gain appears only for well-spread feature distributions, and the worst case remains rank-free. Spectrally, rank sets the price of adaptation: canceling the leading singular directions of the pretrained weight requires both sufficient rank and sufficient budget. We bound the smallest rank achieving a desired source--target alignment, with upper and lower bounds that match under two-sided spectral decay. Together, these results recast rank as governing which updates are reachable and what cancellation costs---not how much capacity the model has.