新LoRA技能应只读而不写
New LoRA Skills Should Read but Never Write
- Shanghai Jiao Tong University(上海交通大学)
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- Jinan University(暨南大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出READ方法,通过规范分解和单向耦合解决LoRA适配器组合中的干扰问题,使新技能只读旧技能输入子空间,在多个基准上显著提升组合性能。
AI中文摘要:
低秩适配器(LoRA)使得针对每个任务微调大型语言模型的成本变得低廉,但将多个独立训练的适配器组合成一个模型仍然困难:在权重空间中合并更新会导致干扰,在所有任务数据上重新训练成本高昂,而在单独适配器之间进行路由则放弃了单一组合模型的目标。我们将困难追溯到每个组合方法都隐式做出的两个选择。LoRA更新允许无限多个等价分解;当适配器单独使用时,这些分解之间的选择是不可见的,但它决定了适配器之间学习到的交互能看到什么。旧技能与新技能之间的耦合同样可以指向任一方向,而方向决定了旧技能是否继续计算它们之前计算的内容。我们引入了READ(适配器增量的只读扩展),它修正了这两个选择:每个适配器被重写为平衡的规范形式,精确保留其更新;耦合仅沿一个方向增长,因此新技能可以读取旧技能的输入子空间,但不能写入其输出子空间。每次追加时唯一可训练的对象是新技能在耦合矩阵中的一行,组合后的更新折叠进基础权重,无推理成本、路由或任务特定规则。我们在四个基准套件和两个模型家族上评估了READ,逐一添加技能。在多个家族中,READ在每个套件上的平均性能都优于由相同适配器构建的最强已发表基线——在SuperGLUE上提高超过20个百分点,在领域套件上提高超过7个百分点——并且几乎所有完整的添加序列最终都高于每个直接基线。因子坐标和耦合方向(单独适配器从不暴露)决定了组合技能能否存活。
英文摘要:
Low-rank adapters (LoRA) make it cheap to fine-tune a large language model once per task, but combining several independently trained adapters into one model remains difficult: merging the updates in weight space causes interference, retraining on all task data is expensive, and routing between separate adapters gives up the goal of a single combined model. We trace the difficulty to two choices that every composition method makes implicitly. A LoRA update admits infinitely many equivalent factorizations; the choice among them is invisible while an adapter serves alone, but it determines what a learned interaction between adapters can see. A coupling between an old skill and a new one can likewise point in either direction, and the direction decides whether the old skills keep computing what they computed before. We introduce READ (Read-only Expansion of Adapter Deltas), which fixes both choices: each adapter is rewritten into a balanced canonical form that preserves its update exactly, and the coupling grows in one direction only, so a new skill can read the input subspaces of old skills but cannot write into their output subspaces. The only trainable object at each append is the new skill's row of the coupling matrix, and the composed update folds into the base weights with no inference cost, routing, or task-specific rules. We evaluate READ across four benchmark suites and two model families, adding skills one at a time. Across several families, READ improves every suite average over the strongest published baselines built from the same adapters---by more than twenty points on SuperGLUE and more than seven points on the domain suite---and nearly all complete addition sequences end above every direct baseline. Factor coordinates and coupling direction, which a lone adapter never exposes, are what decide whether composed skills survive.