AI 中文总结
研究针对LLM推理的权重量化问题,提出CubicQuant参数化非均匀标量格式,其在不同分布下的重构误差优于现有方案,且具备高效GPU执行潜力。
AI 中文摘要
大语言模型推理的权重量化必须平衡自适应重构层级与足够规则的表示,以适配高效的GPU执行。均匀整数将每个组限制在线性格式中;低比特浮点格式采用固定的指数-尾数结构;而学习型码本虽具灵活性,但解码不规则且需额外元数据。我们提出CubicQuant,这是一种参数化非均匀标量格式,在保留密集整数码流的同时,可在每个权重组内调整重构层级。由2个形状参数和1个尺度指定的单调三次曲线,将均匀间隔的幅值码映射为非均匀层级。该格式系列覆盖1-8比特权重有效载荷,包含对称均匀整数量化作为精确特例,且对于有效载荷宽度B和组大小G,每个权重的有效宽度为B + 64/G比特。我们推导了在均匀、高斯和拉普拉斯分布下的总体失真,构建了连续和感知Dynamic-A8载体的拟合目标,并描述了直接打包权重的GPU执行。对于G=128的有限组,每个分布含15360个样本,W4 CubicQuant相较于最优截断四比特均匀整数量化,在均匀样本上重构RMSE降低3.90%,高斯样本上降低13.49%,拉普拉斯样本上降低28.14%;相较于最优枚举四比特有限浮点格式,上述降幅分别为3.90%、9.44%和6.27%。初步H200内核测量显示存在依赖工作负载的交叉点:窄GEMV下模型数据类型执行更快,而随着行数增加,Dynamic A8变得更具优势。这些结果确立了该格式的表示潜力和直接可执行性;下游模型质量和跨设备端到端性能仍是待评估的问题。
英文摘要
Weight quantization for large-language-model inference must balance adaptive reconstruction levels with representations regular enough for efficient GPU execution. Uniform integers constrain each group to a linear grid. Low-bit floating-point formats use a fixed exponent-mantissa structure, while learned codebooks gain flexibility at the cost of irregular decoding and additional metadata. We introduce CubicQuant, a parametric non-uniform scalar format that preserves a dense integer code stream while adapting reconstruction levels within each weight group. A monotonic cubic curve, specified by two shape parameters and one scale, maps uniformly spaced magnitude codes to non-uniform levels. The family spans 1-8-bit weight payloads, contains symmetric uniform integer quantization as an exact special case, and has effective width B + 64/G bits per weight for payload width B and group size G. We derive population distortion under Uniform, Gaussian, and Laplace distributions, formulate continuous and Dynamic-A8-carrier-aware fitting objectives, and describe direct packed-weight GPU execution. For finite groups of G=128 with 15,360 samples per distribution, W4 CubicQuant reduced reconstruction RMSE relative to optimally clipped four-bit uniform integer quantization by 3.90% on Uniform, 13.49% on Gaussian, and 28.14% on Laplace samples. Relative to the best enumerated four-bit finite floating-point format, the reductions were 3.90%, 9.44%, and 6.27%. Preliminary H200 kernel measurements show a workload-dependent crossover: model-dtype execution is faster for narrow GEMV, while Dynamic A8 becomes favorable as row count grows. The results establish the format's representational promise and direct executability; downstream model quality and cross-device end-to-end performance remain open evaluation questions.
Comments23 pages, 1 figure. Technical report