arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解析水蛭格:用于2比特大语言模型权重的融合多壳解码与显存布局

Unfolding the Leech Lattice: Fused Multi-Shell Decoding and VRAM Layouts for 2-Bit LLM Weights

Pier-Jean Malandrino

arXiv 2609.02652首次发表:更新:

发表机构

Scub(Scub)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文实现了水蛭格多壳解码器,设计了GPU显存布局,对比了不同比特布局的性能,在4B、8B、14B模型上实现端到端加速,2比特模型在2.60GB内存中达87.0 tok/s,质量代价随模型尺寸缩小。

AI 中文摘要

水蛭格(Leech lattice)矢量量化在其自身评估协议下具有已报道的最强2比特性能。其核心是解码单个壳层;我们发现该速率所需的多壳解码器尚无实现方案。本文提供了该解码器的实现,并测量了批量大小为1时解码阶段矩阵向量乘法(GEMV)的服务成本。首先,针对完整的301类码本设计了服务路径:离线扩展为GPU布局,以及融合的反量化加矩阵向量乘法内核,该内核在无 warp 分歧的情况下读取布局,并与f64基准进行了验证。其次,显存内(in-VRAM)速率是与磁盘内(on-disk)速率不同的设计维度。在同一进程中对四种比特精确布局进行了计时,结果显示在恒定带宽下(每权重4.80比特,为FP16的2.15倍),二进制比特平面在大小和速度上优于独热(one-hot)掩码。低于4.3比特时会出现第二条不规则流;在3.6比特时,解码不再是移位和掩码操作。第三,部署的4比特(AWQ)和2比特(QTIP)GEMV内核在同一进程中运行。网格(trellis)内核读取的字节数比我们的服务布局少2.40倍,且在其字节限制的相近比例下运行速度快2.27倍:时间差距与流量差距相关,这是展开大到无法放入查找表的码本的代价。第四,有效性边界:网格内核比无权重对照更快,因此我们的启动几何设置了该下限,且在第二种内存层次结构中,每个格臂都低于FP16。在所有臂的输出头相同的情况下,内核与格式路径在4B、8B和14B模型上分别实现了1.11倍、1.29倍和1.41倍的端到端加速;在输出头为int8时,服务后的4B模型在2.60GB内存中达到87.0 tok/s。质量代价方面,4B模型的困惑度为1.38倍,MMLU分数下降14.7点,且在测量的三个模型尺寸中呈缩小趋势。

英文摘要

Leech-lattice vector quantization holds the strongest reported 2-bit quality under its own evaluation protocol. Its kernel decodes one shell; we found no implementation of the multi-shell decoder the rate requires. This paper supplies one and measures its serving cost for decode-phase GEMV at batch 1. First, a serving path for the full 301-class codebook: an offline expansion into GPU layouts and a fused dequantize-plus-matvec kernel reading them without warp divergence, verified against f64. Second, the in-VRAM rate is a design axis distinct from the on-disk rate. Four bit-exact layouts timed in one process show binary bit planes beating one-hot masks on size and speed at constant bandwidth (4.80 bits per weight, 2.15x FP16). Below 4.3 bits a second, irregular stream enters; at 3.6 the decode stops being shifts and masks. Third, deployed four-bit (AWQ) and two-bit (QTIP) GEMV kernels run in the same process. The trellis kernel reads 2.40x fewer bytes than our served layout and runs 2.27x faster at near-equal fractions of their byte bounds: the time gap tracks the traffic gap, the price of unfolding a codebook too large for a lookup table. Fourth, the validity envelope: the trellis kernel outruns our no-weights control, so our launch geometry sets that floor, and on a second memory hierarchy every lattice arm falls below FP16. With the output head held identical across arms, the kernel-and-format path gains 1.11x, 1.29x and 1.41x end to end at 4B, 8B and 14B; with an int8 output head the served 4B reaches 87.0 tok/s in 2.60 GB. The quality cost, 1.38x perplexity and 14.7 MMLU points at 4B, shrinks across the three sizes measured.

Comments21 pages, 5 figures. Preprint, not peer reviewed. Also deposited at Zenodo, doi:10.5281/zenodo.22133606

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑