arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31560cs.LG

OPTQ的泛化行为与正则化的作用

Generalization behavior of OPTQ and the role of regularization

Erin George, Rayan Saab

首次发表
浏览论文内容

中文总结 AI 辅助

本研究分析OPTQ量化算法的泛化误差,证明其与校准误差的关系及随机OPTQ的误差界,并基于正则化项λ提出新选择建议,实验验证其优越性。

中文摘要 AI 辅助

大型神经网络可以通过将权重舍入或“量化”为能用更少比特表示的数值来进行压缩。一种量化算法OPTQ逐步量化神经网络的权重,使得在指定校准数据集上的平方量化误差尽可能小。我们在泛化设置下研究OPTQ及其变体算法随机OPTQ的性能,并推导出当测试点从固定分布中抽取时算法产生的期望平方误差的界。我们证明了两个结果。一个结果将泛化误差与校准数据集上的误差联系起来,该数据集由与测试分布相同分布的独立样本组成。另一个结果对所有足够好的分布(无论校准数据集如何)限定了随机OPTQ的泛化误差。在这两个结果中,正则化项λ都起着重要作用。我们利用这些结果的见解为λ的选择提出了新的建议,并看到与文献中先前的建议相比,这种λ的选择在实验中表现良好。

英文摘要

Large neural networks can be compressed by rounding or "quantizing" their weights to numbers that admit representations with fewer bits. One algorithm for quantization, OPTQ, progressively quantizes the weights of a neural network so that the squared quantization error on a specified calibration dataset is as small as possible. We study the performance of OPTQ and a variant algorithm, stochastic OPTQ, in a generalization setting and derive bounds for the expected squared error accrued by the algorithm when a test point is drawn from a fixed distribution. We prove two results. One result relates the generalization error to the error on a calibration dataset comprising independent samples from the same distribution as the test distribution. The other result bounds the generalization error of stochastic OPTQ for all sufficiently nice distributions, regardless of the calibration dataset. In both of these results, the regularization term $λ$ plays an important role. We use insights from these results to make a new recommendation for the choice of $λ$ and see that this choice of $λ$ preforms favorably in experiments when compared to prior recommendations in the literature.

发表机构

  • University of California, San Diego(加州大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

↑