发表机构
Fujitsu Limited; Institute of Science Tokyo; RIKEN Center for AIP(富士通株式会社; 东京科学大学; 理化学研究所革新智能统合研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出CASA方法,用激活感知指标替代标量代理求解混合精度量化,并通过跨层局部搜索优化位宽,在超低位宽下显著降低困惑度,提升零样本准确率。
AI 中文摘要
混合精度权重量化通常被建模为多选背包问题(MCKP),但现有求解器依赖标量敏感性代理,将每个权重矩阵的Hessian矩阵压缩为单一数值,并独立处理每个模块。我们证明,即使最优的标量代理相对于完整的激活感知二次型也会产生高达$\sqrt{\kappa(\mathbf{A})\kappa(\mathbf{B})}$的乘法失真,其中$\kappa(\mathbf{A})$和$\kappa(\mathbf{B})$分别表示输入侧和输出侧Hessian因子的条件数。对于典型的大语言模型模块,该界限在$10^1$到$10^{13}$之间变化,使得跨模块敏感性排序不可靠。为解决这些局限性,我们提出跨层激活感知敏感性分配(CASA),一种两阶段方法。在第一阶段,标量代理被由Kronecker分解Hessian导出的激活感知指标取代,将MCKP简化为连续松弛具有闭式解的形式。在第二阶段,跨层感知的局部搜索使用端到端模型损失评估位宽更新。在多个大语言模型和不同位宽预算上的实验表明,CASA相比最新的标量代理基线实现了更低的困惑度,尤其是在超低位宽(每权重<3位)下。此外,零样本准确率的提升与每个模型模块上的平均条件数相关,证实了失真界限作为标量代理失效的实际指标。
英文摘要
Mixed-precision weight quantization is commonly formulated as a Multiple-Choice Knapsack Problem (MCKP), yet existing solvers rely on scalar sensitivity proxies that collapse each weight matrix's Hessian into a single number and treat every module independently. We prove that even the optimal scalar proxy incurs multiplicative distortion up to $\sqrt{κ(\mathbf{A})κ(\mathbf{B})}$ relative to the full activation-aware quadratic, where $κ(\mathbf{A})$ and $κ(\mathbf{B})$ denote the condition numbers of the input- and output-side Hessian factors. This bound varies from $10^1$ to $10^{13}$ for typical LLM modules, making inter-module sensitivity ranking unreliable. To address these limitations, we propose Cross-layer Activation-aware Sensitivity Allocation (CASA), a two-phase method. In Stage 1, the scalar proxy is replaced by an activation-aware metric derived from the Kronecker-factored Hessian, reducing the MCKP to a form whose continuous relaxation admits a closed-form solution. In Stage 2, a cross-layer-aware local search evaluates bit-width updates using the end-to-end model loss. Experiments on multiple LLMs across different bit budgets show that CASA achieves lower perplexity than the latest scalar-proxy baselines, especially at ultra-low bit-widths ($<3$ bits per weight). Moreover, the performance gain in zero-shot accuracy tracks the per-model average condition-number over modules, confirming the distortion bound as a practical indicator of scalar-proxy failure.
Comments33 pages, 7 figures