arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

熵约束自适应随机量化

Entropy-Constrained Adaptive Stochastic Quantization

Ran Ben Basat, Yaniv Ben-Itzhak, Michael Mitzenmacher, Shay Vargaftik

arXiv 2608.18147首次发表:更新:

发表机构

University College London; VMware Research by Broadcom; Harvard University(伦敦大学学院; 博通旗下VMware研究院; 哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对自适应随机量化未考虑后续熵编码的精度损失问题,提出ECASQ方法,给出最优与GPU友好近似动态规划方案,实验表明其精度接近最优且速度优势显著。

AI 中文摘要

自适应随机量化(ASQ)是一种新近提出的量化方法,可在保持无偏性的同时针对给定输入优化均方误差(MSE),旨在缓解现代数据与机器学习工作负载的通信和内存瓶颈,包括模型、梯度、KV缓存压缩及最近邻搜索等场景。实际系统可随后用无损熵编码器对量化数据进行压缩,但现有无偏方法(包括ASQ)在选择量化值时未考虑后续编码阶段,导致精度损失。本文构建熵约束自适应随机量化(ECASQ)问题,联合选择自适应量化值以在熵预算和无偏性约束下最小化MSE。针对长度为d的向量及最多s个量化值,本文给出时间复杂度为O(sd²)、空间复杂度为O(d²)的最优动态规划方案,以及时间复杂度为O(sd²)、空间复杂度为O(d)的GPU友好型近似动态规划方案;该近似方案保证解的MSE不超过每条目少用1比特熵的最优解。此外,本文为近似解提供迭代优化流程,实验显示其在获得近最优结果的同时,相比最优解求解器保持显著速度优势。

英文摘要

Adaptive stochastic quantization (ASQ) is a recently introduced quantization approach that optimizes the Mean Squared Error (MSE) for a given input while preserving unbiasedness. It is designed to alleviate the communication and memory bottlenecks of modern data and machine learning workloads, including model, gradient, and KV-cache compression and nearest-neighbor search. Further, practical systems can then compress quantized data with a lossless entropy encoder. However, existing unbiased methods, including ASQ, choose their quantization values without considering this later encoding stage, leaving accuracy on the table. We formulate the Entropy Constrained Adaptive Stochastic Quantization (ECASQ) problem, which jointly selects adaptive quantization values to minimize MSE under an entropy budget and an unbiasedness constraint. We give an optimal dynamic program with $O(sd^2)$ time and $O(d^2)$ space for a length-d vector and at most s quantization values, and a GPU-friendly approximate dynamic program with $O(sd^2)$ time and $O(d)$ space. The approximation guarantees that the solution has an MSE no larger than the optimal solution that uses one fewer bit of entropy per entry. We also provide an iterative refinement procedure for the approximation solution that, in our experiments, yields near-optimal results while retaining a substantial speed advantage over our solver for the optimal solution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑