arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28003cs.LGcs.AI

质量退化约束下性能最大化的大语言模型量化中层级位宽分配方法

A Method for Layer Bit-Width Allocation in LLM Quantization via Performance Maximization Under a Quality-Degradation Constraint

Artem Safronov

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对Gemma-3-1B提出质量退化约束下性能最大化的层级位宽分配方法,结合SA-PTQ敏感性分布优化TensorRT-LLM量化,在RTX 5090上实现最高19.1%延迟降低,为LLM量化提供新方案。

中文摘要 AI 辅助

本文针对Gemma-3-1B提出一种层级位宽分配方法,将问题建模为在退化预算约束(允许的生成质量损失水平)下的性能最大化(延迟降低)问题。该方法不同于文献中耗时耗资源的均匀层级量化方法(如GPTQ或AWQ),也不同于无经证实的性能加速效果的分配方法(如MixLLM或TorchAO)。本文将前期工作SA-PTQ得到的层级敏感性分布,通过激活直通模式应用于TensorRT-LLM中。根据前期步骤引入的分组(5+5、10+10、all26),按块为各层单独确定精度,区分前馈网络(FFN)、注意力机制(Attention)和语言模型头(lm_head)对整体加速的贡献。在RTX 5090上测量了13种W8A8变体的时钟速度,发现对于FFN和lm_head,量化/反量化的时间成本可通过整数运算得到补偿;而对于短上下文长度,Attention则相反:额外的量化步骤会减慢执行速度。由于导出失败且lm_head不可用,本文提出了适用于TensorRT-LLM的SmoothQuant手动实现方案。在同时考虑所有三个标准且退化最小的情况下,找到的最佳方案是FFN 5+5搭配lm_head,实现了11.0%的延迟降低,且质量损失可忽略不计(Top-1一致性为98.90%,困惑度退化+0.85%)。若接受FFN all26搭配lm_head的可接受质量损失,可实现最高19.1%的加速。本文还提出了进一步的优化方向:INT8下的融合注意力内核、KV缓存量化、使用FP8替代INT8,以及类似FFN的部分注意力量化。

英文摘要

This paper proposes a layer bit allocation method for Gemma-3-1B, formulating the problem as performance maximization (latency decrease) given a degradation budget constraint (allowable level of generation quality loss). This approach is different from time- and resource-consuming uniform layer quantization methods that are used in the literature (like GPTQ or AWQ) or allocation methods without proven performance-accelerating effect (like MixLLM or TorchAO). The layer sensitivity profile resulting from our prior work SA-PTQ is applied using the activation pass-through mode inside TensorRT-LLM. For each layer precision is determined individually in blocks, according to a grouping introduced in the prior step (5+5, 10+10, all26), differentiating the contribution of FFN, Attention, and lm_head to the overall speedup. The clock speed was measured for 13 W8A8 variants on an RTX 5090. We find that for FFN and lm_head the time cost of quantization/dequantization is compensated for by the use of integer arithmetic, while for short context lengths, the opposite holds true for Attention: an additional step of quantization slows execution down. We propose a manual implementation of SmoothQuant for TensorRT-LLM which was necessary due to export failures, unavailable for lm_head. The best solution found under joint consideration of all three criteria with minimal degradation was FFN 5+5 with lm_head, providing an 11.0% reduction in latency with negligible quality loss (98.90% Top-1 agreement, +0.85% perplexity degradation). With acceptable quality loss for FFN all26 + lm_head, a speedup up to 19.1% was found possible. We suggest further optimizations: fused attention kernels in INT8, KV-cache quantization, using FP8 instead of INT8 and partial Attention quantization analogous to FFN.

补充信息

↑