低比特循环状态在混合语言模型中的应用
Low-Bit Recurrent States in Hybrid Language Models
浏览论文内容
中文总结 AI 辅助
本文提出一种无需校准数据或训练的混合精度量化方法,利用可观测性格拉姆矩阵和归一化状态范围分配比特,对数量化衰减率,在四位下显著降低负对数似然,六位时接近全精度性能。
中文摘要 AI 辅助
混合语言模型维持固定大小的循环状态,但现有的量化器通常使用八位或更多位数。根据通道衰减率,量化误差持续存在。我们从可观测性格拉姆矩阵推导出失真权重,并将其与归一化状态范围结合,用于混合精度比特分配,无需校准数据、旋转或训练。我们还对衰减率进行对数量化。通过每令牌状态量化,与三种混合模型中的七个基线中最佳者相比,四位平均载荷将多余负对数似然降低了3.3至27.9倍;元数据成本有所变化。在六位时,负对数似然与FP32状态基线相差不到0.005纳特。消融研究将增益分别归因于可变比特宽度、衰减权重和范围归一化。随着写回频率降低,增益减少并依赖于模型和预算。
英文摘要
Hybrid language models maintain fixed-size recurrent states, but existing quantizers typically use eight bits or more. Quantization errors persist according to channel decay rates. We derive distortion weights from the observability Gramian and combine them with normalized state ranges for mixed-precision bit allocation, without calibration data, rotation, or training. We also quantize decay rates logarithmically. With per-token state quantization, a four-bit mean payload reduces excess negative log-likelihood by factors of 3.3--27.9 relative to the best of seven baselines across three hybrid models; metadata costs vary. At six bits, negative log-likelihood differs from the FP32-state baseline by less than 0.005 nats. Ablations separate gains from variable bit widths, decay weighting, and range normalization. With less frequent write-backs, gains diminish and depend on the model and budget.