发表机构
KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对LoRA适配器合并中系数不匹配问题,提出CT-Merging算法,通过平均任务子空间投影仪估计共识方向并分配任务级系数缩放,利用跨任务SVD子空间构建公共基础,在基准测试中比现有方法准确率更高。
AI 中文摘要
LoRA适配器为许多下游任务专门化预训练模型提供了一种有效方法,但每个任务部署一个适配器在推理时需要适配器存储和任务选择。模型合并通过将独立训练的适配器组合成一个多任务适配器来解决此问题。最近基于奇异值分解(SVD)的LoRA合并方法主要关注构建共享或特定任务方向,而分配给最终方向的系数通常直接来自原始任务SVD。在固定合并基础上,继承的系数保持具有高秩相关性的分量顺序,但其大小与任务更新诱导的系数有很大差异。为解决这种不匹配,我们提出CT-Merging,一种LoRA感知合并算法,从平均任务子空间投影仪估计共识方向并在最终更新中分配任务级均方根系数缩放。CT-Merging利用跨任务SVD子空间的重复支持来构建公共基础,同时在方向构建后减少对按秩SVD大小的依赖。在DC-Merge CLIP适配器基准测试中,CT-Merging与现有最先进合并方法相比实现了更高的平均归一化准确率,在ViT-B/32上比DC-Merge进一步提高2.56个点,在ViT-L/14 KnoTS训练的检查点上提高1.51个点。
英文摘要
LoRA merging methods increasingly operate on the low-rank structure of task updates, yet how the common subspace is estimated and how coefficients are assigned after recomposition are rarely compared directly. We propose CT-Merging, which estimates common directions from averaged task subspace projectors and assigns a separate residual scale to each task. Projector averaging selects directions supported across task subspaces without weighting them by singular magnitude, while task-specific scaling removes component-wise magnitude variation and preserves scale differences across tasks. On the released KnOTS CLIP adapters, CT-Merging achieves the best average and worst-task normalized accuracy on both backbones, improving over the strongest baseline by up to 2.56 and 6.65 points, respectively. On the DC-Merge adapter benchmark, it achieves the best average normalized accuracy in eight of nine backbone and task-count settings. Ablations show that projector averaging outperforms summed-update SVD and that task-specific scaling improves worst-task accuracy over global isotropic scaling.
Comments5 pages, 1 figure