arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LoRA需要多少秩?Transformer注意力的秩-误差边界

How Much Rank Does LoRA Need? Rank-Error Bounds for Transformer Attention

Gerard Conangla Planes

arXiv 2608.26052首次发表:更新:

发表机构

Aily Labs(Aily Labs)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对Transformer注意力,建立了LoRA秩与近似误差的理论边界,分析了秩的选择规律,扩展了相关分析至融合多头LoRA及联合查询/键更新的场景。

AI 中文摘要

选择低秩适配(LoRA)更新的秩通常是一项经验性任务。本文针对Transformer注意力,提供了一种与任务相关的理论,用于确定每个LoRA秩可实现的近似误差。我们固定一个预训练的注意力头、一个目标注意力函数,以及下游任务输入的分布,对秩为r的查询LoRA更新可实现的最小期望Kullback-Leibler(KL)误差进行边界限定。当目标注意力概率被限定为远离零时,我们证明误差的下界与ψ(∥d∥₂)成比例,其中d是候选注意力分数与目标注意力分数之间的差值,且ψ(t)=min{t²,t}。我们还证明了无条件上界为min{∥d∥₂²/4,√2∥d∥₂}。在明确的可实现性、几何和矩条件下,我们随后将最佳秩为r的误差限定在ψ(√Tᵣ)的明确倍数与min{Tᵣ/4,√{2Tᵣ}}之间,其中Tᵣ是目标更新的下游加权尾部能量。当候选分数保持在目标分数的固定范围内时,我们还提供了目标Fisher边界;当一部分token承载大部分概率质量时,我们提供了无限制的下界。这些谱边界描述了有限分数近似。随后,我们构建了明确的族,其中softmax饱和使得匹配注意力函数所需的秩严格小于匹配有限logits所需的秩。最后,我们将分析扩展到融合多头LoRA以及联合查询/键更新,揭示了秩共享和查询/键分解约束的影响。

英文摘要

Choosing the rank of a low-rank adaptation (LoRA) update is usually an empirical task. In this paper, we provide a task-dependent theory of the approximation error achievable at each LoRA rank for Transformer attention. We fix a pretrained attention head, a target attention function, and a distribution over inputs from the downstream task, and bound the smallest expected Kullback--Leibler (KL) error achievable by a rank-$r$ query LoRA update. When target attention probabilities are bounded away from zero, we prove a lower bound of the error proportional to $ψ(\|d\|_2)$, where $d$ is the difference between candidate and target attention scores and $ψ(t)=\min\{t^2,t\}$. We also prove an unconditional upper bound $\min\{\|d\|_2^2/4,\sqrt2\|d\|_2\}$. Under explicit realizability, geometry, and moment conditions, we then bound the best rank-$r$ error between an explicit multiple of $ψ(\sqrt{T_r})$ and $\min\{T_r/4,\sqrt{2T_r}\}$, where $T_r$ is the downstream-weighted tail energy of the target update. We also provide target-Fisher bounds when candidate scores remain within a fixed range of the target scores, and an unrestricted lower bound when a subset of tokens carries most of the probability mass. These spectral bounds describe finite-score approximation. We then construct explicit families in which softmax saturation makes the rank required to match the attention function strictly smaller than the rank required to match the finite logits. Finally, we extend the analysis to fused multi-head LoRA and joint query/key updates, exposing the effects of rank sharing and query/key factorization constraints.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑