RoSA:用于内存高效微调的旋转稀疏自适应
RoSA: Rotational Sparse Adaptation for Memory-Efficient Fine-Tuning
浏览论文内容
中文总结 AI 辅助
RoSA通过旋转稀疏适配层子集,减少内存占用,同时保持微调性能,可与PEFT结合。
中文摘要 AI 辅助
参数高效微调(PEFT)通过将训练聚焦于一小部分参数子集来降低适配基础模型的成本。与此思路互补,我们引入了RoSA(旋转稀疏自适应),它将适配范围缩小到每次仅适配一部分层。RoSA在整个训练过程中冻结靠近输入的较低层,并在较后的层上旋转一个可训练块,逐步增加靠近输入的冻结层数量。这种设计减少了优化器状态内存,缩短了反向传播,甚至如果缓存了最后一个冻结层的激活,还能缩短前向传播。由于RoSA与可训练参数化的选择正交,它可以与PEFT方法或稀疏优化器在每个活动块内结合使用。跨多个LLM架构和任务的实验表明,RoSA在保持强大微调性能的同时降低了峰值内存。
英文摘要
Parameter-efficient fine-tuning (PEFT) reduces the cost of adapting foundation models by focusing training on a small parameter subset. Complementary to this idea, we introduce RoSA (Rotational Sparse Adaptation), which narrows adaptation to a subset of layers at a time. RoSA freezes lower layers close to the input throughout training and rotates a trainable block over later layers, progressively increasing the number of frozen layers close to the input. This design reduces optimizer-state memory, shortens backpropagation, and even forward propagation if activations at the last frozen layer are cached. Because RoSA is orthogonal to the choice of trainable parameterization, it can be combined with PEFT methods or sparse optimizers within each active block. Experiments across multiple LLM architectures and tasks show that RoSA reduces peak memory while maintaining strong fine-tuning performance.
发表机构
- Saarland University(萨尔兰大学)
- CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心)
机构由 AI 辅助整理,请以论文原文为准。