TORQUE:在旋转前后优化量化(或不量化)的内容
TORQUE: Optimizing What (not) to Quantize Before and After Rotation
浏览论文内容
中文总结 AI 辅助
TORQUE通过联合优化旋转前后高精度保留的坐标数量与选择,在固定比特预算下改进随机旋转量化,实现更优的重建精度与存储成本权衡。
中文摘要 AI 辅助
均匀随机旋转是量化的一种有效预处理步骤:它们使归一化坐标分布近似高斯分布,从而能够使用离线优化的码本。我们引入了TORQUE,一个框架,通过联合优化在旋转前后以高精度保留多少坐标以及保留哪些坐标,在固定的总体期望比特预算下,改进了先前使用随机旋转的量化工作。直观上,在旋转前,以高精度保留大的输入坐标可以通过防止旋转将其值分散到多个坐标上来减少总体误差。同样,在旋转后,以高精度保留一小部分最大幅度的坐标,使得剩余值能够使用针对截断高斯分布离线优化的码本进行更精确的量化。我们推导了一个量化误差上界,并证明了对于每个k,旋转前保留top-k坐标可以最小化该上界。这将坐标子集的搜索简化为对k的优化,使得快速优化器能够使用离线码本和并行参数选择进行实际实现。我们通过在高斯模型下的数值评估以及最近邻检索、KV缓存压缩和激活压缩实验,展示了重建精度与存储成本之间改进的权衡。
英文摘要
Uniform random rotations are an effective preprocessing step for quantization: they make normalized coordinate distributions approximately Gaussian, enabling the use of codebooks optimized offline. We introduce TORQUE, a framework that improves on previous quantization works that use random rotations by jointly optimizing how many and which coordinates to preserve at high precision both before and after rotation, under a fixed overall expected bit budget. Intuitively, before rotation, preserving large input coordinates at high precision can reduce overall error by preventing the rotation from spreading their values across many coordinates. Likewise, after rotation, preserving a small fraction of the largest-magnitude coordinates at high precision allows the remaining values to be quantized more accurately using codebooks optimized offline for the resulting truncated Gaussian distribution. We derive a quantization error upper bound and prove that top-$k$ pre-rotation retention minimizes it for each $k$. This reduces the search over coordinate subsets to an optimization over $k$, enabling a fast optimizer that uses offline codebooks and parallel parameter selection for practical implementation. We demonstrate an improved tradeoff between reconstruction accuracy and storage cost through numerical evaluation under the Gaussian model and experiments on nearest-neighbor retrieval, KV-cache compression, and activation compression.
发表机构
- University College London(伦敦大学学院)
- Harvard University(哈佛大学)
- VMware Research by Broadcom(博通旗下威睿研究院)
机构由 AI 辅助整理,请以论文原文为准。