发表机构
Henselian; University of Copenhagen; Towards Deep Learning(Henselian; 哥本哈根大学; Towards Deep Learning)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ApexQuant 是一种免校准的量化方法,通过递归再量化残差误差并利用随机旋转实现再各向同性化,在无数据条件下达到接近全精度的性能,尤其以二比特结果最优。
AI 中文摘要
我们引入了 ApexQuant,一种免校准的量化方法,它递归地对残差误差进行再量化,作为现有量化器之上的精炼层。我们证明了新的随机旋转会将每个残差返回到超球面上的均匀分布,这表征了连续遍历中渐进误差衰减的速率。这一结果使我们能够在读取任何权重之前,确定一个层为达到目标权重空间误差所需的遍历次数。每个前缀本身就是一个有效的低比特率模型,因此一个工件可以服务于多种精度。我们用三个可互换的阶段实例化 ApexQuant:标量、$E_8$ 和网格,并在四个开放权重的大型语言模型以及地球观测和医学领域上进行了验证,在这些领域中,由于限制性许可或患者材料受隐私约束,分布内数据通常无法获得。渐进再各向同性化在四比特时接近全精度的几个百分点之内,并在完全无数据设置下提供了我们测量到的最佳二比特结果。
英文摘要
We introduce ApexQuant, a calibration-free quantization method that recursively re-quantizes the residual error, serving as a refinement layer on top of existing quantizers. We establish that a fresh random rotation returns each residual to the uniform distribution on the hypersphere, which characterizes the rate of progressive error decay across successive passes. This result lets us determine, before any weight is read, how many passes a layer needs for a target weight-space error. Every prefix is itself a valid lower-rate model, so one artifact serves several precisions. We instantiate ApexQuant with three interchangeable stages, scalar, $E_8$ and trellis, and validate it on four open-weight LLMs and on Earth-observation and medical domains where in-distribution data is often unattainable as imagery arrives under restrictive licences or due to patient material under privacy constraints. Progressive re-isotropization comes within a few percent of full precision at four bits and gives the best two-bit arm we measure, in a completely data-free setting.
Comments25 pages, 6 figures