StagQ:面向LLM的约束驱动多精度权重量化
StagQ: Constraint-Driven Multi-Precision Weight Quantization for LLMs
浏览论文内容
中文总结 AI 辅助
StagQ提出一种多精度权重量化格式,通过共享仿射映射和稀疏侧记录,在多个精度下以更高速率实现优于基线的LLM性能。
中文摘要 AI 辅助
在多个部署环境中服务大型语言模型(LLM)需要多个权重精度工作点。多精度格式通过一个流来服务所有这些精度,该流的前缀是有效的低精度代码,而不是存储多个副本。我们提出了StagQ,一种多精度权重格式,其主流是一个2位分组仿射基,随后在二进阶梯调度上配置数量可变的1位细化平面。每个支持的精度都是一个可读的前缀,通过从所有精度共享的元数据导出的仿射映射进行解码,无需逐权重查找。一个稀疏的侧记录,在网格拟合前后均被填充,用于保留网格服务最差的少数权重。我们报告了编码器的两种配置。在2位时,较便宜的配置在Llama-3.1-8B、Phi-4和OLMo-2-7B上比最强的多精度基线领先3.1到7.0个MMLU点,同时逻辑速率略低。在3位时,它在Llama-3.1-8B上领先,在Phi-4上以更高速率领先,并在OLMo-2-7B上持平。在4位时,它在所有三个模型上以更高速率持平。在NVIDIA A100 GPU上的批量一矩阵-向量乘积中,使用合成权重计时,我们的内核在大多数形状-精度情况下比两个基线内核更快。
英文摘要
Serving a large language model (LLM) across a fleet of deployments requires several weight-precision operating points. Multi-precision formats serve them all from one stream whose prefixes are valid lower-precision codes, instead of storing multiple copies. We present StagQ, a multi-precision weight format whose main stream is a 2-bit group-wise affine base followed by a configurable number of 1-bit refinement planes on a dyadic step schedule. Every supported precision is a readable prefix, decoded by an affine map derived from metadata shared across all precisions, with no per-weight lookup. A sparse side record, filled both before and after the grid is fitted, holds out the few weights the grid serves worst. We report two configurations of the encoder. At two bits the cheaper one leads the strongest multi-precision baseline on Llama-3.1-8B, Phi-4, and OLMo-2-7B by 3.1 to 7.0 MMLU points, at a slightly lower logical rate. At three bits it leads on Llama-3.1-8B, leads on Phi-4 at a higher rate, and ties on OLMo-2-7B. At four bits it ties on all three, at a higher rate. In a batch-one matrix-vector product on an NVIDIA A100 GPU, timed on synthetic weights, our kernel is faster than the two baseline kernels in most shape-precision cases.
发表机构
- Huawei Technologies Ltd.(华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。