发表机构
KAUST; ISTA; Yandex Research(沙特阿卜杜拉国王科技大学; 奥地利科学技术研究所; Yandex 研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大型语言模型微调成本高的问题,提出Super和Supra方法,利用剪枝显著性信号,结合Wanda分数与LoRA,通过预算分配规则保持参数预算,在算术实验中取得高准确率,证明简单排序可为PEFT提供支持。
AI 中文摘要
大型语言模型(LLMs)的微调成本高昂,因为全参数更新需要大量内存、计算资源和每个任务的存储空间。我们研究了最初为剪枝开发的显著性信号是否可用于选择模型适应的位置。我们提出了Super,一种稀疏参数高效微调(PEFT)方法,它使用从校准过程计算出的Wanda风格激活加权幅度分数来固定一个小的可训练支持集。然后我们引入了Supra,一种混合适配器,它将这种稀疏更新与LoRA相结合,同时通过简单的预算分配规则保持匹配的可训练参数预算。在对Llama-3.2-1B和Meta-Llama-3-8B进行的单种子Math17K算术实验中,最佳的Super/Supra变体在测试的调度选择适配器配置中实现了最高的平均准确率。我们还包括一个仅基于幅度的PaFi风格支持作为最接近的无训练稀疏基线,并发现幅度和Wanda风格排序下的低分支持集可能是有效的。这些结果表明,简单的剪枝启发式排序可为PEFT提供有用的固定稀疏支持,特别是与低秩适配器结合时。
英文摘要
Large language models (LLMs) remain expensive to fine-tune because full-parameter updates require substantial memory, compute, and per-task storage. We study whether saliency signals originally developed for pruning can be reused to choose where a model should adapt. We propose Super, a sparse parameter-efficient fine-tuning (PEFT) method that fixes a small trainable support using a Wanda-style activation-weighted magnitude score [Sun et al., 2023] computed from a calibration pass. We then introduce Supra, a hybrid adapter that combines this sparse update with LoRA while preserving a matched trainable-parameter budget through a simple budget-splitting rule. In single-seed Math17K arithmetic experiments on Llama-3.2-1B and Meta-Llama-3-8B, the best Super/Supra variants achieve the highest average accuracy among the tested schedule-selected adapter configurations. We also include a PaFi-style magnitude-only support as a closest training-free sparse baseline and find that low-score supports under both magnitude and Wanda-style orderings can be effective. These results suggest that simple pruning-inspired orderings can provide useful fixed sparse supports for PEFT, especially when combined with low-rank adapters.
Comments26 pages, 3 figures, 19 tables. Code: https://github.com/vectozavr/SuperTuning