发表机构
Scub(Scub)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Tetra提出基于李奇格点的新码本,将大语言模型量化至每参数约2.7比特,通过高效解码核函数减少内存读取,在保持性能的同时实现高吞吐生成。
AI 中文摘要
李奇格点量化在每权重两比特下具有良好的质量,但其码本包含超过10^14个点,对于查找表而言数量过多。我们之前的核函数在加载时展开码字,并为2比特的码字从GPU内存读取每权重4.804比特。我们提出Tetra,一种基于同一格点的新码本。一个24权重块仍占用48比特,其中大部分索引一个64状态的Golay码网格和一个共享的16 KiB表。该核函数在矩阵-向量乘积内部通过六次表加载和两次小型查找来解码一个块,并读取每权重2.148比特。对于完整模型,我们为每矩阵行重新训练一个缩放因子,将损失最大的矩阵存储为4比特整数,并用4比特嵌入表来补偿。我们的Qwen3-4B、8B和14B文件在整个模型上分别保持每参数2.73、2.70和2.73比特。它们在完整MMLU测试集上得分分别为63.37、69.58和75.66,比每参数5.3至6.0比特的4比特AWQ低4.76、4.21和2.46分。在我们的引擎中,它们每秒生成113.8、95.0和57.2个token。在GSM8K上,通过所服务的核函数,它们相对于FP16损失9.63、4.62和3.26分。在4B规模下,我们的文件比此http URL的IQ2_XXS(每参数2.48比特)高出23.6分。我们为表格或图表测量的每个数字均来自一块NVIDIA L40S GPU。我们预先注册了主要实验。
英文摘要
Leech-lattice quantization gives good quality at two bits per weight, but its codebooks hold more than 10^14 points, too many for a lookup table. Our earlier kernel expanded the codes at load time and read 4.804 bits per weight from GPU memory for 2 bits of code. We present Tetra, a new codebook on the same lattice. A 24-weight block still takes 48 bits, most of which index a 64-state trellis of the Golay code and one shared 16 KiB table. The kernel decodes a block with six table loads and two small lookups inside the matrix-vector product, and reads 2.148 bits per weight. For full models, we retrain one scale per matrix row, store the matrices that lose the most as 4-bit integers, and pay for them with 4-bit embedding tables. Our Qwen3-4B, 8B and 14B files hold 2.73, 2.70 and 2.73 bits per parameter over the whole model. They score 63.37, 69.58 and 75.66 on the full MMLU test set, 4.76, 4.21 and 2.46 points below 4-bit AWQ at 5.3 to 6.0 bits per parameter. They generate 113.8, 95.0 and 57.2 tokens per second in our engine. On GSM8K, through the served kernel, they lose 9.63, 4.62 and 3.26 points to FP16. At 4B our file scores 23.6 points above llama.cpp's IQ2_XXS (2.48 bits per parameter). Every number we measured for a table or figure comes from one NVIDIA L40S GPU. We preregistered the main experiments.
Comments15 pages, 4 figures, 8 tables. Code, measurement logs and preregistrations: https://github.com/pjmalandrino/llvq