arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

冻结核心需要任务信号:用于低资源大语言模型适配的Fisher白化交叉协方差

Frozen Cores Need Task Signal: Fisher-Whitened Cross-Covariance for Low-Resource LLM Adaptation

Wentao Ye, Zhanming Shen, Zhiqing Xiao, Yao Ding, Haobo Wang, Gang Chen

arXiv 2609.00762首次发表:更新:

发表机构

Zhejiang University; Hunyuan, Tencent(浙江大学; 腾讯混元)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出FCCA方法,通过固定权重矩阵基并优化小尺寸核心,在低资源预算下实现大语言模型高效适配,在多任务多模型上表现优异,参数规模远低于LoRA、DoRA等方法。

AI 中文摘要

参数高效微调通常被视为更新多少参数的问题,但在可训练状态预算极为有限的情况下,这些系数的作用位置同样重要。我们通过冻结核心适配研究这一选择:校准阶段为每个权重矩阵固定左右基,微调仅优化一个r×r的核心。这消除了可训练因子修复不良初始张成的能力,使子空间质量可直接观测。我们提出FCCA,该方法估计带符号的输入-误差交叉协方差,用对角Fisher矩对其白化,在所得局部度量中截断,将选定方向映射回,并应用薄QR分解以获得稳定的核心坐标。在匹配的r²预算下,我们在11个任务、4种模型设置和3个随机种子上比较8种基构造器。在Qwen2.5-3B上,FCCA达到83.0的宏平均,比下一个最佳匹配预算构造器高2.3个点,且在全部11个任务上均优于其未白化的RawGrad对照组;在所有3种Qwen规模上均排名第一,在Llama-3.2-1B上与最佳方法的差距在0.13个点以内。受控消融实验显示,白化带来2.7至17.2个点的增益,并确定QR是测试场景中稳定核心优化的必要条件。最后,FCCA在优化3.69万个参数(而非约740万个)的情况下,与LoRA和DoRA的平均差距分别为0.32和0.23个点。这些结果表明,精心选择的固定张成可在可训练参数和优化器状态成本低得多的情况下,恢复大部分可移动低秩因子的优势。

英文摘要

Parameter-efficient fine-tuning is usually framed as a question of how many parameters to update. Under a severe trainable-state budget, however, where those coefficients act is equally consequential. We study this choice through frozen-core adaptation: a calibration pass fixes left and right bases for each weight matrix, and fine-tuning optimizes only an $r\times r$ core. This removes the ability of trainable factors to repair a poor initial span and makes subspace quality directly observable. We introduce FCCA, which estimates the signed input--error cross-covariance, whitens it with diagonal Fisher moments, truncates it in the resulting local metric, maps the selected directions back, and applies thin QR to obtain stable core coordinates. Under a matched $r^2$ budget, we compare eight basis constructors on 11 tasks, four model settings, and three seeds. On Qwen2.5-3B, FCCA reaches an 83.0 macro-average, 2.3 points above the next-best matched-budget constructor, and exceeds its unwhitened RawGrad control on all 11 tasks. It ranks first at all three Qwen scales and finishes within 0.13 points of the best method on Llama-3.2-1B. Controlled ablations show gains of 2.7--17.2 points from whitening and identify QR as necessary for stable core optimization in the tested regime. Finally, FCCA comes within 0.32 and 0.23 average points of LoRA and DoRA while optimizing 36.9K rather than roughly 7.4M parameters. These results show that a carefully selected fixed span can recover most of the benefit of movable low-rank factors at a much smaller trainable and optimizer-state cost.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑