发表机构
Royal Holloway, University of London; University of West London(伦敦大学皇家霍洛威学院; 西伦敦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出LARA,一种在冻结模型残差流中运行的轻量适配方法,在参数数量相同时性能与LoRA相当,支持多种行为平滑插值与按token路由,可在单个模型上高效托管多种行为。
AI 中文摘要
我们提出了LARA(轻量加法残差适配),这是一种高效适配方法,它在冻结模型的残差流中运行,而非在其权重中运行。LoRA是向权重矩阵添加低秩更新,而LARA则读取少量层的隐藏状态,并向残差流添加低秩校正,保留所有基础权重不变。在代码微调任务和偏好优化(DPO)上,参数数量相同时LARA的性能与LoRA相当。由于适配是冻结基础模型加上残差,LARA在推理时会应用一个缩放因子γ,可在基础行为与适配行为之间平滑插值,这种分级控制是权重空间适配所不具备的特性。此外,由于每种行为都是共享冻结基础之上的小型残差模块,因此可以同时驻留多种行为,并按token自动路由。我们在一个冻结的15亿参数模型上放置了7种行为,其中6种为微调得到,1种为偏好优化得到,总开销约为33MB,而若每种行为单独使用完整模型则会占用更多资源。由于基础模型未被改动,各行为可独立训练,且按token选择而非按需加载,这适合在单个设备的单个模型上托管多种行为并添加新行为。
英文摘要
We present LARA (Lightweight Additive Residual Adaptation), a method for efficient adaptation that operates in the residual stream of a frozen model rather than in its weights. Where LoRA adds an update of low rank to weight matrices, LARA reads the hidden state at a small set of layers and adds a correction of low rank back to the residual stream, leaving all base weights untouched. On a code fine-tuning task and on preference optimization (DPO), LARA matches LoRA at equal parameter counts. Because adaptation is a frozen base plus a residual, LARA exposes a scale γ, applied at inference, that interpolates smoothly between base and adapted behavior, a form of graded control that adaptation in weight space does not offer. Finally, because each behavior is a small residual module over a shared frozen base, many behaviors can be held resident at once and routed automatically per token. We place seven behaviors, six fine-tuned and one optimized for preference, on one frozen 1.5B model for roughly 33 MB of overhead, against one full model for each behavior. Because the base is untouched, behaviors are trained separately and selected per token rather than loaded on demand, which suits hosting many behaviors, and adding new ones, on a single model on a device.