arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20988cs.LGcs.AI

面向大语言模型量化鲁棒性的雅可比引导噪声注入

Jacobian-guided Noise Injection for Quantization Robustness in Large Language Models

发表机构亚马逊
查看机构详情
  • Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

Deepanshu Pandey, Arnav Chavan, Nahush Lele, Sankalp Dayal, Deepak Gupta

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对大语言模型量化时自注意力机制对离散化误差敏感的问题,提出雅可比引导噪声注入训练策略,在低比特量化下显著提升了模型在图像分类与语言建模任务中的性能。

中文摘要 AI 辅助

大语言模型(LLMs)的量化常因自注意力机制对离散化误差敏感而受阻。我们发现softmax算子是量化稳定性的瓶颈,因其对异常值及状态依赖的雅可比敏感。我们从理论上证明,抑制该雅可比的范数有助于限制量化引发的性能下降。基于此,我们提出雅可比引导噪声注入(Jacobian-Guided Noise Injection),这是一种训练策略,在注意力前logits中注入零均值高斯噪声,其方差直接由雅可比Frobenius范数推导。与依赖启发式方法或直接惩罚雅可比的现有方法不同,我们的方法提供了一种基于局部注意力敏感性确定最优噪声方差的途径。我们在SOTA大语言模型架构上评估该方法,结果显示其相较于流行的PTQ方法具有更高的鲁棒性。实证分析表明,该方法在低比特量化设置下,针对SigLIP模型在ImageNet-1K数据集上的Top-1准确率可实现高达37%的相对提升,针对语言模型在WikiText数据集上的相对困惑度可实现高达40%的改善,证明了该方法的有效性。

英文摘要

Quantization of Large Language Models (LLMs) is often hindered by the sensitivity of the self-attention mechanism to discretization errors. We identify the softmax operator as a bottleneck for quantization stability due to its sensitivity to outliers and state-dependent Jacobian. We theoretically establish that suppressing the norm of this Jacobian helps in bounding quantization-induced performance degradation. Based on this, we propose Jacobian-Guided Noise Injection, a training strategy that injects zero-mean Gaussian noise into pre-attention logits, with variance derived directly from the Jacobian Frobenius norm. Unlike prior approaches that rely on heuristic or penalise jacobian directly, our method provides a way to identify the optimal noise variance based on the local attention sensitivity. We evaluate the method on SOTA LLM architectures, where it demonstrates improved robustness over popular PTQ methods. Empirical analysis reveals that the proposed method gives up to +37% relative gains on Top-1 accuracy on ImageNet-1K for SigLIP and improves relative perplexity by upto 40% on WikiText for language models in low bit quantisation settings, proving the efficacy of the approach.

补充信息

↑