REAL-Q:基于动态梯度下降的端到端大语言模型量化方法
REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent
浏览论文内容
中文总结 AI 辅助
针对现有大语言模型训练后量化方法的信息失配问题,REAL-Q采用端到端损失对齐代理与动态块级梯度下降,在主流模型W4A16量化下将端到端KL散度最高降低约49%。
中文摘要 AI 辅助
训练后量化(PTQ)是在严格资源约束下部署大语言模型(LLM)的关键技术。当前最先进的PTQ方法采用单一闭式二阶求解器对每层进行量化:为保持解析可处理性,它们会大幅近似全局损失(舍弃跨通道耦合、将输出行池化为组),随后冻结整个层的海森矩阵,无法在损失景观逐列变化时刷新——我们将这一现象称为信息失配。我们提出REAL-Q(实时端到端损失对齐大语言模型量化,Real-time E2E-loss Aligned LLM Quantization),一种新型PTQ范式,打破了这种权衡:不再为解析可处理性弱化目标,REAL-Q以全局损失的端到端对齐代理为目标,在每128列组成的列块后应用细粒度动态块级梯度下降对其优化。通过将这种细粒度校正与用于平滑跨层过渡的滑动窗口机制结合,REAL-Q有效缓解了网络中的误差传播。在LLaMA-3.1(8B和70B)及Qwen3(0.6B-32B)模型的W4A16量化设置下,REAL-Q相比最先进的全局引导方法,将端到端KL散度降低了约49%。
英文摘要
Post-training quantization (PTQ) is essential for deploying large language models (LLMs) under strict resource constraints. State-of-the-art PTQ methods quantize each layer with a single closed-form second-order solver: to remain analytically tractable, they heavily approximate the global loss (dropping cross-channel coupling, pooling output rows into groups), and they then freeze the resulting Hessian across the entire layer, with no way to refresh it as the loss landscape shifts column by column--a phenomenon we call information misalignment. We propose REAL-Q (Real-time E2E-loss Aligned LLM Quantization), a novel PTQ paradigm that breaks this compromise: instead of diluting the objective for the sake of analytic tractability, REAL-Q targets an end-to-end-aligned surrogate of the global loss and refines it via fine-grained, dynamic Block-wise Gradient Descent applied after every column block (128 columns). By coupling this fine-grained correction with a sliding window mechanism for smooth cross-layer transitions, REAL-Q effectively mitigates error propagation across the network. On LLaMA-3.1 (8B and 70B) and Qwen3 (0.6B-32B) at W4A16, REAL-Q reduces end-to-end KL divergence by up to ~49% relative to state-of-the-art globally-guided methods.
发表机构
- Peking University(北京大学)
- Northeastern University(东北大学)
- ZTE Corporation(中兴通讯股份有限公司)
机构由 AI 辅助整理,请以论文原文为准。