arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AQLoRA:一种用于快速量化LoRA微调的零搜索方案

AQLoRA: A Zero-Search Recipe for Fast Quantized LoRA Fine-Tuning

Md Romyull Islam

arXiv 2608.23816首次发表:更新:

发表机构

Kennesaw State University(肯尼索州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AQLoRA是无需搜索、无需校准的快速量化LoRA微调方案,通过保留部分层为fp16跳过反量化提升速度,在多模型数据集上验证了其速度与准确率表现,还总结了共享硬件的计时规则。

AI 中文摘要

量化微调(QLoRA)可节省内存但无法缩短时间,它会实时对每一个4位权重进行反量化,因此其训练速度比fp16 LoRA更慢。我们提出了AQLoRA(自适应量化LoRA),这一方案可挽回部分时间开销。仅需对权重进行一次CPU遍历即可完成所有设置,无需搜索也无需校准数据。该遍历会根据NF4重建误差对层进行排名,并在内存预算内将前K层保留为fp16格式,这些层可跳过反量化,这正是速度提升的来源。其中,质量设置可适配每一层,速度设置仅适配顶部块,因此反向传播可提前停止。该规则能在数秒内精确复现Unsloth手动整理的动态4位选择,而基于搜索的分配则需要多次校准遍历。我们在Commonsense-170K数据集上对6个模型和4个架构族(规模从14亿到140亿参数)进行了评估。速度设置的训练速度比调优良好的QLoRA快11.1±2.7%,并损失约1个准确率点;在9次独立计时会话中均表现更快,最慢时快7%。质量设置的训练速度快4.8±2.4%,其准确率在所有模型上与QLoRA持平,且与fp16 LoRA的准确率相差在1个点以内,仅多占用0.2 GiB内存。这些误差条是在独立会话之间测量的,而非单个会话内。获得这些结果的过程中,我们总结了共享硬件计时的三条规则:固定测量时长而非步数;通过重复分支测量本底噪声,而非近似方法;重复整个会话,单次扫描内计算的本底噪声会低估数倍的真实不确定性,且随机种子几乎无法控制不确定性。我们通过对照实验验证了该方案,并报告了两个失败的对照:按权重密度选择适配器层的效果不优于随机选择;按量化误差选择受保护层的效果也不佳。受保护层的数量而非其身份决定了速度效果。

英文摘要

Quantized fine-tuning (QLoRA) saves memory but not time. It dequantizes every 4-bit weight on the fly, so it trains more slowly than fp16 LoRA. We present AQLoRA (Adaptive-Quantization LoRA), a recipe that buys part of that time back. One CPU pass over the weights sets everything, with no search and no calibration data. The pass ranks layers by NF4 reconstruction error and keeps the top-K in fp16 under a memory budget. Those layers skip dequantization, which is where the speed comes from. A quality setting adapts every layer. A speed setting adapts only the top blocks, so the backward pass stops early. The rule reproduces Unsloth's hand-curated dynamic-4bit selection exactly, in seconds, where search-based allocation needs repeated calibration passes. We evaluate on Commonsense-170K across six models and four architecture families, from 1.4B to 14B. The speed setting trains 11.1 +/- 2.7% faster than well-tuned QLoRA and gives up about one accuracy point. It was faster in all nine independent timing sessions, at worst by 7%. The quality setting trains 4.8 +/- 2.4% faster. Its accuracy is level with QLoRA on every model and within a point of fp16 LoRA, for 0.2 GiB more memory. These error bars are measured between independent sessions, not within one. Earning them taught us three rules for timing on shared hardware. Fix the measurement duration, not the step count. Measure the noise floor from a duplicated arm, not a nearly identical method. Repeat whole sessions: a floor computed inside one sweep understates the real uncertainty several times over, and the random seed controls almost none of it. We validate the recipe with controls and report the two that failed. Choosing adapter layers by weight density is no better than random. Choosing protected layers by quantization error is not either. The count of protected layers, not their identity, carries the speed effect.

Comments23 pages, 5 figures, 19 tables. Code and experimental artifacts: https://github.com/Romyull-Islam/AQLoRA

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑