发表机构
University of Wisconsin-Madison; Google DeepMind(威斯康星大学麦迪逊分校; 谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出FrameFT方法,采用融合帧基的稀疏系数矩阵实现Transformer模型的参数高效微调,性能优于或媲美现有PEFT技术,且所需可训练参数更少。
AI 中文摘要
参数高效微调(PEFT)策略如低秩适配(LoRA)是对大规模预训练模型进行微调的有效方案;然而其内存需求随模型规模扩展,为$\boldsymbol{O}(dr)$,其中$d$为模型的隐藏维度,$r$为秩。我们提出的FrameFT方法,用融合帧(Fusion Frame)基中的稀疏系数矩阵对参数更新$\boldsymbol{\triangle W}$进行建模。融合帧可通过算法生成并在模型各层间共享,从而实现极为高效的更新。仅需存储和优化基展开的稀疏系数,即可减少内存占用。FrameFT中系数矩阵的稀疏结构以及融合帧的稀疏性,带来了显著的计算优势,且我们的分析提供了形式化收敛结果。我们在一组监督微调基准上对该思路进行了评估,重点关注语言任务,同时也报告了其在视觉模型上的应用。实验表明,FrameFT的性能与最先进的PEFT技术相当或更优,但所需的可训练参数少得多。
英文摘要
Parameter-Efficient Fine-Tuning (PEFT) strategies such as Low-Rank Adaptation (LoRA) are effective solutions for fine-tuning large-scale pre-trained models; however, their memory requirements scale with the size of the model, $\mathcal{O}(dr)$, where $d$ is the model's hidden dimension and $r$ is the rank. Our proposal, FrameFT, models the parameter update $ΔW$ with a sparse coefficient matrix in a Fusion Frame basis. Fusion Frames can be generated algorithmically and shared across model layers, enabling very efficient updates. Only the sparse coefficients of the basis expansion are stored/optimized, reducing the memory footprint. The sparse structure of the coefficient matrix in FrameFT and the sparsity in the Fusion Frames give large compute benefits, and our analysis provides formal convergence results. We evaluate the idea across a suite of supervised fine-tuning benchmarks, focusing on language tasks, but also report application to vision models. Our experiments show that FrameFT achieves performance on par with/exceeding state-of-the-art PEFT techniques, but needs far fewer trainable parameters.
Comments21 pages, 6 figures