arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2407.08044cs.CLcs.AIcs.LG

RoLoRA:微调旋转的无异常值LLM以实现有效的权重-激活量化

RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization

  • Hong Kong University of Science and Technology(香港科技大学)
  • Meta Reality Labs(元宇宙现实实验室)

机构由 AI 辅助整理,请以论文原文为准。

Xijie Huang, Zechun Liu, Shih-Yang Liu, Kwang-Ting Cheng

更新

AI总结:

针对LoRA流程中权重-激活量化因激活异常值导致性能下降的问题,提出首个基于旋转消除异常值并保持无异常值特性的LoRA量化方案RoLoRA,显著提升了低比特量化的收敛性与鲁棒性。

AI中文摘要:

低秩微调(LoRA)作为参数高效微调(PEFT)的代表性方法,通过仅更新大型语言模型(LLM)中一小部分权重,显著提升了训练效率。近期,仅权重量化技术也被应用于LoRA方法以减少微调的内存占用。然而,将权重-激活量化应用于LoRA流程仍待深入探索,且我们观察到由于激活异常值的存在,性能出现显著下降。在本研究中,我们提出了RoLoRA,这是首个用于有效权重-激活量化的基于LoRA的方案。RoLoRA利用旋转消除异常值,并提出旋转感知微调以在旋转的LLM中保持无异常值特性。实验结果表明,RoLoRA在权重-激活设置下持续改善了低比特LoRA收敛性和训练后量化鲁棒性。我们在LLaMA2-7B/13B、LLaMA3-8B模型上评估了RoLoRA,与LoRA基线相比,在常识推理任务中,4比特权重-激活量化的LLaMA2-13B实现了高达29.5%的绝对准确率提升。我们进一步在大型多模态模型(LLaVA-1.5-7B)上证明了其有效性。代码可在https://github.com/HuangOwen/RoLoRA获取。

英文摘要:

Low-Rank Adaptation (LoRA), as a representative Parameter-Efficient Fine-Tuning (PEFT)method, significantly enhances the training efficiency by updating only a small portion of the weights in Large Language Models (LLMs). Recently, weight-only quantization techniques have also been applied to LoRA methods to reduce the memory footprint of fine-tuning. However, applying weight-activation quantization to the LoRA pipeline is under-explored, and we observe substantial performance degradation primarily due to the presence of activation outliers. In this work, we propose RoLoRA, the first LoRA-based scheme for effective weight-activation quantization. RoLoRA utilizes rotation for outlier elimination and proposes rotation-aware fine-tuning to preserve the outlier-free characteristics in rotated LLMs. Experimental results show RoLoRA consistently improves low-bit LoRA convergence and post-training quantization robustness in weight-activation settings. We evaluate RoLoRA across LLaMA2-7B/13B, LLaMA3-8B models, achieving up to 29.5% absolute accuracy gain of 4-bit weight-activation quantized LLaMA2- 13B on commonsense reasoning tasks compared to LoRA baseline. We further demonstrate its effectiveness on Large Multimodal Models (LLaVA-1.5-7B). Codes are available at https://github.com/HuangOwen/RoLoRA

补充信息

↑