arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30908cs.LG

在量化代码空间中利用梯度微调低比特模型

Fine-Tuning Low-Bit Models with Gradient in Quantized Code Space

Shiguang Wu, Zhouchen Lin, Quanming Yao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对低比特模型微调的效率与部署适配问题,提出GradCodes方法,通过代码代理梯度加速优化并以引导搜索保证部署适配,在多任务上实现低比特模型微调性能提升。

中文摘要 AI 辅助

微调低比特模型旨在适配量化模型,同时保持最终部署的检查点为相同的低比特形式,该设置在实践中十分重要,因为它能降低存储和部署的内存与推理成本。在这一约束下,适配成为了针对量化代码和缩放因子的优化问题。现有的连续低比特训练效率较高,但可能因直通估计误差或后量化间隙而产生失真;离散搜索符合部署要求,但在有限的训练预算下往往效率过低。我们提出代码代理梯度作为可部署代码空间中的一阶信号以加速优化,并执行引导搜索以保持部署适配性。在算术推理、指令跟随和结构化语言理解任务上的实验表明,GradCodes(代码代理梯度与引导搜索结合的方法)在不同量化数据类型下均能持续改进低比特模型的微调效果。代码提供在此处的链接。

英文摘要

Fine-tuning Low-bit models aims to adapt a quantized model while keeping the final deployed checkpoint in the same low-bit form. This setting is practically important as it reduces memory and inference cost for storage and deployment. Under this constraint, adaptation becomes an optimization problem over quantization codes and scales. Existing continuous low-bit training is efficient, but it can be distorted by straight through estimation error or by post-quantize gap; discrete search is deployment-faithful, but it is often too inefficient under a finite training budget. We propose code surrogate gradient as the first order signal in deployable code space to acceleate optimization, and performing guided search to preserve deployment faithfulness. Experiments across arithmetic reasoning, instruction following, and structured language understanding show that GradCodes consistently improves fine-tuning low-bit models across different quantization datatypes. Code is provided at https://github.com/ovo67/GradCodes.

发表机构

  • Tsinghua University(清华大学)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

↑