arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不确定信号是否有用?对带有回滚机制的不确定感知解码的系统研究

Do Uncertainty Signals Help? A Systematic Study of Uncertainty-Aware Decoding with Rollback Mechanisms

Xianzong Wu, Xiaohong Li, Yuejun Guo, Xinyang Liu, Tianlin Li, Junjie Wang, Qiang Hu

arXiv 2608.14653首次发表:更新:

AI 中文总结

本文通过系统研究带有回滚机制的不确定感知解码,发现其可提升代码LLM的生成性能,其中token熵等信息论不确定信号效果最优,反馈引导的回滚是主要改进来源。

AI 中文摘要

预测不确定性是一种被广泛采用的用于量化模型置信度的指标,其下游应用涵盖模型解释、数据选择以及预测回滚。尽管其效用已得到证实,但不确定性量化在增强大型语言模型(LLM)代码生成方面的潜力仍在很大程度上未被探索,这引发了一个关键问题:不确定性在多大程度上可以作为改善基于LLM的代码生成的有效信号?为回答该问题,本文研究了不确定感知回滚解码,这是一种推理时策略,利用不确定信号识别不可靠的生成区域,并回滚到更早的有效前缀,无需重新训练模型。我们在统一的解码设置下,对7种代码LLM、5个代码生成基准以及8种 token 级不确定性信号评估了该框架。结果显示,完整回滚框架在评估的基准和模型设置上优于等预算重启,在功能性代码生成基准上,pass@1提升最高达0.26,AvgTestPassRate提升最高达0.35;在Dsec-Python上,Patch-Aligned Safe Rate的绝对提升最高达6.4%。在评估的信号中,token熵、负对数似然等信息论指标呈现出最有利的整体趋势,在标准基准上常取得最佳或接近最佳的结果。组件控制的消融实验进一步表明,反馈引导的回滚是主要的改进来源,而在检查、预算、回滚和分支衰减固定的情况下,不确定性定位会提供额外增益。

英文摘要

Prediction uncertainty is a widely adopted metric for quantifying model confidence, with downstream applications spanning model explanation, data selection, and prediction rollback. Despite its demonstrated utility, the potential of uncertainty quantification to enhance code generation in large language models (LLMs) remains largely underexplored, raising a critical question: to what extent can uncertainty serve as an effective signal for improving LLM-based code generation? To answer this question, we study uncertainty-aware rollback decoding, an inference-time strategy that uses uncertainty signals to identify unreliable generation regions and roll back to earlier valid prefixes without retraining the model. We evaluate this framework on seven code LLMs, five code generation benchmarks, and eight token-level uncertainty signals under a unified decoding setup. Our results show that the complete rollback framework improves over equal-budget restart across the evaluated benchmarks and model settings, with gains of up to 0.26 in pass@1 and 0.35 in AvgTestPassRate on functional code generation benchmarks, and an absolute improvement of up to 6.4\% in Patch-Aligned Safe Rate on Dsec-Python. Among the evaluated signals, information-theoretic measures such as token entropy and negative log-likelihood show the most favorable overall trend, frequently achieving the best or near-best results on standard benchmarks. A component-controlled ablation further shows that feedback-guided rollback provides the main improvement, while uncertainty localization provides an additional gain when checking, budget, rollback, and branch decay are held fixed.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑