发表机构
School of Artificial Intelligence, Jilin University; Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, MOE, China; International Center of Future Science, Jilin University(吉林大学人工智能学院; 教育部知识驱动的人机智能工程研究中心; 吉林大学未来科学国际合作中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对大语言模型中LoRA方法的不足,提出SOS-LoRA,将秩更新重参数化为多个静态低秩专家之和,通过分解总秩、固定多尺度缩放及正交初始化等,实现性能提升且可完全合并,实验证明其优于基线和变体。
AI 中文摘要
低秩自适应(LoRA)是一种广泛用于大语言模型的参数高效微调(PEFT)方法。在固定秩预算下,LoRA通过单个低维输入侧路径对每个适配权重进行参数化,这可能会通过共享输入方向耦合异构行为并在优化过程中引发干扰。我们提出了静态正交子空间LoRA(SOS-LoRA),它是一种直接插入式扩展,将秩为rtot的更新重新参数化为K个静态(始终开启、无路由)低秩专家的总和。SOS-LoRA分解专家之间的总秩,应用固定多尺度缩放方案以鼓励尺度分离的优化动态,并通过跨专家正交初始化和轻量级正则化促进不同的输入侧方向。SOS-LoRA仍然完全可合并,合并后不增加推理时间参数或延迟。在推理和知识密集型基准测试(Llama 2/3)、基于编码器的自然语言理解(GLUE)和数学推理(GSM8K/MATH)上的实验表明,与匹配预算的LoRA基线和最新变体相比,SOS-LoRA有持续的性能提升。代码可在这个https网址获取。
英文摘要
Low-Rank Adaptation (LoRA) is a widely used parameter-efficient fine-tuning (PEFT) method for large language models. Under a fixed rank budget, LoRA parameterizes each adapted weight through a single low-dimensional input-side pathway, which may couple heterogeneous behaviors through shared input directions and induce interference during optimization. We propose Static Orthogonal Subspace LoRA (SOS-LoRA), a drop-in extension that reparameterizes a rank-rtot update as a sum of K static (always-on, non-routed) low-rank experts. SOS-LoRA (i) decomposes the total rank across experts, (ii) applies a fixed multi-scale scaling scheme to encourage scale-separated optimization dynamics, and (iii) promotes diverse input-side directions via cross-expert orthogonal initialization and a lightweight regularizer. SOS-LoRA remains fully mergeable, adding no inference-time parameters or latency after merging. Experiments on reasoning and knowledge-intensive benchmarks (Llama 2/3), encoder-based NLU (GLUE), and math reasoning (GSM8K/MATH) show consistent gains over matched-budget LoRA baselines and recent variants. Code is available at https://github.com/llm172/sos-lora.