发表机构
School of Computer Science and Technology, Soochow University; Hithink Research; Electronic Engineering, Tsinghua University(苏州大学计算机科学与技术学院; 海思睿研究中心; 清华大学电子工程系)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM部署的资源瓶颈,该研究提出FAMPWQ方法,通过费舍尔信息度量层敏感度结合强化学习分配位宽,在7个模型和5个基准上优于7种基线方法,提升了量化性能。
AI 中文摘要
近年来,大语言模型(LLMs)在多个领域取得了显著成就,但其过高的资源需求阻碍了在资源受限设备上的部署。尽管模型量化是一种有效的方法,但传统量化方法通常因采用均匀位宽或简单启发式的敏感度评估而导致严重的性能下降。本文提出了一种新颖的基于费舍尔信息的自适应混合精度权重量化方法,即FAMPWQ,用于在商用GPU上实现高效的LLM推理。首先,我们提出了一个系统模型,采用新颖的费舍尔信息度量来衡量各层对量化的敏感度。其次,我们在FAMPWQ中提出了一种基于强化学习的位宽分配器,该分配器基于费舍尔信息敏感度度量生成自适应位宽分配策略。在7个模型和5个基准上进行的大量实验表明,FAMPWQ在困惑度(PPL)上显著优于7种基线方法,其PPL最多可降低3.39;在准确率上最多可提高6.87%;在LLM-as-a-judge对比中,其胜率最高可达76%。
英文摘要
Recent years have witnessed remarkable achievements of Large Language Models (LLMs) in multiple domains, while the excessive resource requirements of LLMs hinder the deployment on resource-constrained devices. Although model quantization stands out as an effective approach, conventional quantization approaches typically incur severe performance degradation due to uniform bit-width or simple heuristic sensitivity evaluation. In this paper, we propose a novel Fisher information-based Adaptive Mixed Precision Weight Quantization approach, i.e., FAMPWQ, which performs layer-adaptive weight quantization for effective LLM inference on commodity GPUs. First, we propose a system model with a novel Fisher information metric to measure the layer-wise sensitivity to quantization. Second, we propose a reinforcement learning-based bit-width allocator in FAMPWQ, which generates an adaptive bit-width allocation strategy based on the Fisher information sensitivity metric. Extensive experiments on 7 models and 5 benchmarks demonstrate that FAMPWQ significantly outperforms 7 baseline approaches in terms of PPL (up to 3.39 smaller), accuracy (up to 6.87% higher), and LLM-as-a-judge comparison (up to 76% win rate).
Comments21 pages, to appear in EMNLP 2026