发表机构
Portland State University(波特兰州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对STE在超低位宽QAT中的梯度失配问题,提出Q-MINO最小范数优化器,结合梯度一致性、状态漂移正则化与对齐约束,经Frank-Wolfe求解,理论保证渐近收敛,实验验证有效。
AI 中文摘要
直通估计器(STE)是量化感知训练(QAT)中广泛使用的启发式方法,但其替代梯度可能与底层量化目标存在显著不匹配,导致更新噪声和参数振荡,尤其在超低位宽场景下。我们提出了量化感知最小范数优化器(Q-MINO),一种时间束方法,结合梯度一致性、状态漂移正则化和对齐约束,从近期优化状态构建稳定且最小范数的更新方向。Q-MINO使用带可行回退初始化的热启动Frank-Wolfe过程求解所得约束子问题。理论上,通过随机Lyapunov Kurdyka-Łojasiewicz(KL)框架,我们证明Q-MINO实现渐近邻域收敛。此外,我们详细介绍了Q-MINO在不同量化下的数值实验。
英文摘要
The Straight-Through Estimator (STE) is a widely used heuristic for Quantization-Aware Training (QAT), but its surrogate gradients can exhibit substantial mismatch with the underlying quantized objective, leading to noisy updates and parameter oscillations, particularly in ultra-low-bit regimes. We propose the Quantization-Aware Minimal-Norm Optimizer (Q-MINO), a temporal bundle method that combines gradient consensus, state-drift regularization, and an alignment constraint to construct stabilized, minimum-norm update directions from recent optimization states. Q-MINO solves the resulting constrained subproblem using a warm-started Frank--Wolfe procedure with a feasible fallback initialization. Theoretically, via a stochastic Lyapunov Kurdyka--Łojasiewicz (KL) framework, we show that Q-MINO achieves asymptotic neighborhood convergence. Moreover, we detail numerical experiments with Q-MINO at various quantizations.