arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

极端低位LLM中量化移动的上下文效用

Contextual Utility of Quantization Moves in Extreme Low-Bit LLMs

Wenxuan Xiao, Xu Cao

arXiv 2609.09867首次发表:更新:

发表机构

Astrmira Tech.(Astrmira科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示极端低位LLM中量化移动的效用具有上下文依赖性,提出中点评估与精确端点束搜索方法,以提升量化后模型的准确性与困惑度。

AI 中文摘要

后训练量化器使用重建代理或局部损失近似来选择有限的代码更改,但量化移动的效用取决于执行该移动时所处的状态。我们识别出这种上下文依赖性的两个来源。首先,移动的位移很重要:在移动中点评估梯度可以捕获沿移动累积的曲率,而当前状态的线性化则忽略了这一点。在来自Llama-3.2模型的冻结两位移动中,中点评估预测精确端点损失变化方向比当前状态梯度准确得多。其次,移动相互作用:合法量化状态的穷举格可以很好地由二次伪布尔函数近似,但它们的小成对分量可以决定帕累托前沿,并导致不同的评估泛函偏好相反的方向。这些效应解释了重建最优代码重新选择和加性组合的失败。在各自的中点读取每个移动修复了局部选择步骤,并提高了下游准确性和保留困惑度,而更大的支持需要从实际达到的状态评估精确端点。精确端点束搜索找到的稀疏更改优于更大的单次更新,并且在干预更改后重新定价相同的移动会产生广泛的符号反转。这些结果表明,量化效用在几个移动的粒度上是上下文的:可靠的构建必须沿着它们自己的路径评估有限更改,并从演化的量化状态中组合它们。

英文摘要

Post-training quantizers select finite code changes using reconstruction proxies or local loss approximations, but the utility of a quantization move depends on the state through which it is executed. We identify two sources of this contextual dependence. First, the displacement of the move matters: evaluating the gradient at the move midpoint captures curvature accumulated along the move that a current-state linearization omits. Across frozen two-bit moves from Llama-3.2 models, midpoint evaluation predicts the direction of exact endpoint loss changes substantially more accurately than current-state gradients. Second, moves interact: exhaustive lattices of legal quantized states are well approximated by quadratic pseudo-Boolean functions, yet their small pairwise components can determine Pareto fronts and cause different evaluation functionals to prefer opposite directions. These effects explain failures of reconstruction-optimal code re-selection and additive composition. Reading each move at its own midpoint repairs the local selection step and improves downstream accuracy and held-out perplexity, while larger supports require evaluating exact endpoints from the state actually reached. Exact-endpoint beam search finds sparse changes that dominate much larger one-shot updates, and repricing the same moves after intervening changes produces widespread sign reversals. These results show that quantization utility is contextual at the granularity of a few moves: reliable construction must evaluate finite changes along their own paths and compose them from the evolving quantized state.

CommentsPreprint. 5 figures. Includes appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑