arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于量化推理模型的校准e-CUSUM解码:为何令牌对数概率不适用于解码监测

Calibrated e-CUSUM Decoding for Quantized Reasoning Models: Why Token Log-Probability Is the Wrong Observable for Decoding Monitors

El Hassane Ettifouri, Ayoub Belfatmi, Mahaman Sanoussi Yahaya Alassan, Walid Dahhane

arXiv 2607.11317首次发表:更新:

发表机构

Novelis Research(诺贝丽斯研究公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对低比特量化推理模型思维链退化问题,指出用居中令牌对数概率监测解码不可行。提出结合退化感知警报分数与校准e过程检测器的解码控制器,经GSM8K实验验证其有效性及不足,给出方法层面的解释与替代方案。

AI 中文摘要

低比特量化使小型推理模型部署成本低廉,但会削弱其思维链。这促使在生成变得不可靠时进行干预的解码器端监测器出现。我们表明,一个自然候选指标,即居中令牌对数概率增量$\log p(w_t)+H_t$,在此目的下是错误的观测指标。在模型自身采样法则下,它本质上是均值为零的鞅,所以它衡量的是采样自一致性而非轨迹健康状况,并且在置信重复期间几乎无信号,此时$\log p(w_t)$和熵都接近零。我们引入一种无需训练的解码控制器,它结合了(i)一种融合令牌不确定性与显式逐字重复的退化感知警报分数,以及(ii)一种受校准e过程启发的顺序检测器。原始乘积过程在条件均值为零的情况下是维尔有效的,而部署的基于CUSUM下限的统计量被视为经验变化检测器,因为分数依赖历史且自相关。在GSM8K数据集上,使用FP16和INT4的DeepSeek-R1-Distill-Qwen-1.5B模型进行实验,校准将一个在93%-95%的生成中触发的监测器转变为失败轨迹的选择性检测器($\phi\approx0.3$,精度约为0.6,基础率为0.38)。在这个试点中,控制器减少了测量的逐字退化信号,并使INT4精度从63%提高到69%,虽有正向变化但统计学上无定论(配对麦克内马尔检验$p = 0.18$,$n = 100$),令牌预算成本增加28%。我们还发现,在GSM8K上,非终止而非循环是主要的失败模式。主要贡献在于方法层面:解释了为何居中令牌对数概率不适用于解码器监测,并给出了一个经过校准、谨慎评估的替代方法。

英文摘要

Low-bit quantization makes small reasoning models inexpensive to deploy but can degrade their chains of thought. This motivates decoder-side monitors that intervene when generation becomes unreliable. We show that a natural candidate, the centered token log-probability increment $\log p(w_t)+H_t$, is the wrong observable for this purpose. Under the model's own sampling law it is a mean-zero martingale by construction, so it measures sampling self-consistency rather than trajectory health and is nearly silent during confident repetition, where both $\log p(w_t)$ and entropy are close to zero. We introduce a training-free decoding controller that combines (i) a degeneration-aware alarm score fusing token uncertainty with explicit verbatim repetition and (ii) a calibrated e-process-inspired sequential detector. The raw product process is Ville-valid under a conditional-mean null, while the deployed CUSUM-floored statistic is treated as an empirical change detector because the score is history-dependent and autocorrelated. On GSM8K with DeepSeek-R1-Distill-Qwen-1.5B in FP16 and INT4, calibration turns a monitor that fires on 93--95% of generations into a selective detector of failing traces ($ϕ\approx 0.3$, precision $\approx 0.6$ against a 0.38 base rate). In this pilot, the controller reduces measured verbatim-degeneration signals and yields a positive but statistically inconclusive INT4 accuracy change from 63% to 69% (paired McNemar $p=0.18$, $n=100$), at a 28% token-budget cost. We also find that non-termination, rather than looping, is the dominant failure mode on GSM8K. The main contribution is methodological: an explanation of why centered token log-probability is inadequate for decoder monitoring and a calibrated, cautiously evaluated replacement.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑