arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14659cs.AIcs.LGcs.SE

当不确定性还不够时:代码生成中自校正的实证研究

When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation

Pranav Rakasi, Maanas Lalwani, Arnav Srivastava, Arya Palanivel, Tinuade Adeleke, Ruizhe Li, Sean Wu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过实证发现,针对代码生成的不确定性估计方法难以可靠提升生成准确率,仅基于验证的自校正策略可显著提升Pass@1指标,廉价不确定性估计器仅适合作为校正循环的门控信号。

中文摘要 AI 辅助

用于代码生成的大语言模型常生成错误解决方案却无可靠的失败指示。本研究探讨为自然语言开发的不确定性估计方法是否可迁移至代码生成,以及此类信号能否通过选择性自校正提升代码生成效果。我们在HumanEval和BigCodeBench两个基准上,针对三个小型代码大语言模型,评估了五种不确定性方法:平均token熵、口头置信度、P(True)、集成熵以及语义熵探测。研究发现,多样本P(True)与正确性的相关性最强,而包括语义熵探测在内的其他所有方法仅产生弱相关性。随后,我们利用这些不确定性信号驱动三种自校正策略:自适应解码、基于不确定性的再生以及基于验证的再生。研究结果揭示了一个比预期更强烈的负面发现:基于不确定性的自校正无法可靠提升Pass@1指标,在两个基准的6种配置中有5种配置的准确率出现下降,降幅为3个百分点至10个百分点;自适应解码在6种配置中有4种配置的准确率出现下降。仅基于验证的自校正可可靠提升Pass@1指标,在HumanEval上提升6至26个百分点,在BigCodeBench上提升8至20个百分点,提升幅度与基线模型强度呈负相关。这些发现在两个基准上均保持一致,表明廉价的不确定性估计器本身不足以提升代码正确性,其实际价值在于作为更昂贵的基于执行的校正循环的门控信号,而非作为验证的独立替代品。

英文摘要

Large language models for code generation often produce incorrect solutions without reliable indicators of failure. We study whether uncertainty estimation methods developed for natural language transfer to code generation, and whether such signals can improve code generation via selective self-correction. We evaluate five uncertainty methods: mean token entropy, verbalized confidence, $P(\text{True})$, entropy ensembles, and semantic entropy probes, across three small code LLMs on HumanEval and BigCodeBench. We find that multi-sample $P(\text{True})$ achieves the strongest correlation with correctness, while all the other methods, including semantic entropy probes, yield only weak correlation. We then use these uncertainty signals to drive three self-correction policies: adaptive decoding, uncertainty-based regeneration, and verification-based regeneration. Our results reveal a stronger negative finding than anticipated: uncertainty-based self-correction fails to reliably improve Pass@1, degrading accuracy in 5 of 6 configurations across both benchmarks ($-3$pp to $-10$pp), and adaptive decoding degrades accuracy in 4 of 6 configurations. Only verification-based self-correction reliably improves Pass@1, with gains of $+6$ to $+26$ percentage points on HumanEval and $+8$ to $+20$ percentage points on BigCodeBench, scaling inversely with baseline strength. These findings replicate consistently across both benchmarks and suggest that cheap uncertainty estimators are insufficient on their own to improve code correctness, and that their practical value lies in serving as gating signals for costlier execution-based correction loops rather than as standalone substitutes for verification.

发表机构

  • University of Michigan(密歇根大学)
  • New York University(纽约大学)
  • University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
  • Algoverse AI
  • University of Aberdeen(阿伯丁大学)
  • University of Oxford(牛津大学)

机构由 AI 辅助整理,请以论文原文为准。

↑