arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

规范量化:从语言模型对称性中在线学习量化最优基

GaugeQuant: Online Learning of Quantization-Optimal Bases from LLM Symmetries

Miguel P. Bento, João F. Seabra

arXiv 2607.20757首次发表:更新:

AI 中文总结

研究利用Transformer内部连续对称性,通过在损失中引入LogSumExp项破坏对称性,在线学习量化最优基,无需校准数据和量化模拟,训练开销小,在不同量化设置下降低了LLaMA - 2 7B模型的困惑度。

AI 中文摘要

已知Transformer具有内部连续对称性,其在修改量化时使输出不变。规范量化通过在损失中引入LogSumExp项来利用这种训练中的对称性破坏,从而选择使激活异常值最小化的基。一个停止梯度算子确保仅更新旋转矩阵,使语言建模目标完全不变。该方法无需特定校准数据和量化模拟,训练开销可忽略不计。在W4A4量化、组大小为128的情况下,LLaMA - 2 7B模型的困惑度从8.22降至6.73,与需要冻结模型和校准数据集的训练后方法竞争。在W4A16下,困惑度从11.16降至5.45。代码可在该https网址获取。

英文摘要

Transformers are known to have internal continuous symmetries that leave outputs invariant, while modifying quantization. GaugeQuant leverages this in-training by introducing a LogSumExp term to the loss that breaks the symmetries, thus selecting a basis that minimizes activation outliers. A stop-gradient operator ensures that only rotation matrices are updated, yielding the language modeling objective completely unaltered. Our requires no specific calibration data, no quantization simulation, and adds negligible training overhead. With the LLaMA-2 7B model under W4A4 quantization with group size 128, perplexity drops from 8.22 to 6.73, competing with post-training methods that require frozen models and calibration datasets. Under W4A16, perplexity drops from 11.16 to 5.45. Code is available at https://github.com/MPedraBento/gauge-quant.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑