arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

JIVEAdapter:一种基于联合与个体变异解释(JIVE)的多任务加性低秩适配器

JIVEAdapter: A Multi-Task Additive Low-Rank Adapter via Joint and Individual Variation Explained (JIVE)

Sara Abdali, Pashmina Cameron

arXiv 2610.07036首次发表:更新:

发表机构

Microsoft Applied Sciences Group (ASG)(微软应用科学组)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

JIVEAdapter提出多任务加性低秩适配器,将权重更新分解为共享联合结构与任务特有正交个体结构,自适应分配秩,冻结联合部分重用,在GLUE/SuperGLUE上以匹配秩达到强基线性能。

AI 中文摘要

参数高效微调以远低于全量微调的成本适配预训练模型,然而大多数低秩适配器是单任务的,并以乘法方式表示每次权重更新,未能明确区分跨任务共享的内容与任务特有的内容。我们提出JIVEAdapter,一种受统计学中联合与个体变异解释(JIVE)启发的多任务“加性”低秩适配器。JIVEAdapter将每次权重更新分解为一个跨所有任务共享的联合结构加上一个每任务特有的个体结构,对个体结构施加惩罚使其与联合结构近似正交,从而保持共享与任务特有信号的“可解释性”和分离性,并在共享联合池和每任务个体池之间自适应地分配秩。联合结构仅学习一次,可在任务组上联合学习或增量地逐任务学习,之后被冻结并作为新任务的先验重用,无需重新训练共享部分。在GLUE和SuperGLUE上使用DeBERTaV3-base,JIVEAdapter在匹配的每任务有效秩下与强大的单任务和多任务低秩基线竞争,无需MoE等额外模块;当存在相关的已包含任务时,其冻结的联合结构通过仅使用廉价的每方向缩放重用该任务的个体结构来服务保留任务,否则训练一个小的新个体结构。

英文摘要

Parameter-efficient fine-tuning adapts pretrained models at a fraction of the cost of full fine-tuning, yet most low-rank adapters are single-task and represent each weight update multiplicatively, leaving no explicit account of what is shared across tasks and what is task-specific. We introduce JIVEAdapter, a multi-task "additive" low-rank adapter inspired by statistical Joint and Individual Variation Explained (JIVE). JIVEAdapter decomposes every weight update into a Joint structure shared across all tasks plus a per-task Individual structure, penalizes the Individual structures to be near-orthogonal to the Joint so shared and task-specific signal stay "interpretable" and separated, and allocates rank adaptively across a shared Joint pool and a per-task Individual pool. The Joint is learned once, jointly over a task group or incrementally, one task at a time, then frozen and reused as a prior for new tasks without retraining the shared part. On GLUE and SuperGLUE with DeBERTaV3-base, JIVEAdapter is competitive with strong single-task and multi-task low-rank baselines at a matched per-task effective rank, without extra modules such as MoE, and when a related held-in task exists its frozen Joint serves a held-out task by reusing that task's Individual with only a cheap per-direction scale, otherwise training a small new one.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑