arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型量化的变换:大逆变换与格式协同设计

Transforms for LLM Quantization: The Great Inversion and Format Co-Design

Ehsan Jokar

arXiv 2608.25188首次发表:更新:

AI 中文总结

本研究针对LLM量化的变换阶段,提出大逆变换原则,分析变换与数字格式的协同设计,综述200项研究分类43种变换方法,提炼部署指南并指出开放问题。

AI 中文摘要

目前多数有竞争力的4位大语言模型(LLM)研究流程都遵循相同的起始步骤:应用一种线性、保留函数的变换(旋转、缩放、置换、非正交仿射变换),使异常值更有利于匹配分组缩放因子,之后再进行量化。然而,目前尚无针对该变换阶段的综述,相关文献一直在悄然重新推导一套旧理论。我们识别并形式化了组织该变换的核心原则——大逆变换:分配灵活的编码方案有利于能量集中,而部署的矩阵指令所采用的分组共享缩放因子量化则有利于组内扁平化。经典变换编码(1963年提出:去相关、分配比特、量化)在固定总比特率下为每个坐标分配不同比特;对于高速率下的高斯源,卡胡宁-洛埃夫变换(Karhunen-Loeve transform)的能量集中特性可最小化失真。而部署的操作数块每个分组携带一个绝对最大缩放因子,所有位置使用相同比特,无比特分配;在均匀网格上,该目标有利于扁平化,可通过哈达玛非相干性(Hadamard incoherence)实现。我们证明组内优序(majorization)下存在对立性:两种方案的指向相反,每种方案都有针对自身目标的证明,对于一般频谱,不存在最优性保证可跨方案传递。第二个核心维度是数字格式:非均匀FP4网格使扁平化的收益降低,MXFP4的2的幂次分组缩放因子仍有利于限制在该分组内的旋转,而NVFP4的尾数携带缩放因子则在很大程度上消除了这种收益,因此目标极点同时取决于分配机制和数字格式。我们综述了截至2026年6月的200项研究,按结构、数据感知性、搜索式与构造式、运行时成本对43种变换方法进行分类;记录了这些方法与GPTQ量化的组合效果(若有报告);针对不同部署场景提炼了首选方案指南;最后指出了该领域存在的开放问题。

英文摘要

Most competitive 4-bit LLM research pipelines now open the same way: apply a linear, function-preserving transform (rotation, scaling, permutation, non-orthogonal affine) so the outlier mass sits more favorably against the group scales, and only then round. Yet we are aware of no survey dedicated to this transform stage, and its literature is quietly re-deriving an older theory. We identify and formalize the principle that organizes it, the Great Inversion: allocation-flexible coding rewards energy concentration, whereas the grouped shared-scale quantization a deployed matrix instruction performs rewards within-group flattening. Classical transform coding (1963: decorrelate, allocate bits, quantize) spends different bits per coordinate at a fixed total rate; for a Gaussian source at high rate the Karhunen-Loeve transform's concentration minimizes distortion. A deployed operand tile instead carries one absolute-maximum scale per group and equal bits everywhere, with no allocation; on a uniform grid that objective rewards flattening, approached by Hadamard incoherence. We prove that opposition under within-group majorization: the prescriptions point in opposite directions, each backed by a proof against its own objective, and for a generic spectrum no optimality guarantee transfers. A second axis is the number format: the non-uniform FP4 grid makes flattening buy less, MXFP4's power-of-two block scale still rewards a rotation confined to that block, and NVFP4's mantissa-carrying scale largely removes that pull, so the target pole depends jointly on allocation regime and format. We survey 200 works to a June 2026 cutoff; classify 43 transform methods by structure, data-awareness, searched-versus-constructed, and runtime cost; record, where reported, how they compose with GPTQ rounding; distill a first-choice guide by deployment regime; and close with the open problems it exposes.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑