arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于大语言模型的代码生成中过度自信失败的表征

Characterizing Overconfident Failure in LLM-Based Code Generation

Ravishka Rathnasuriya, Wei Yang

arXiv 2610.11300首次发表:更新:

发表机构

University of Texas at Dallas; Fudan University; Institute of Systems for Advanced Computing(达拉斯大学; 复旦大学; 先进计算系统研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对四个开源代码模型和三个执行基准,研究代码生成LLMs的过度自信失败问题,发现现有不确定性信号仅部分有效,指令调优无法可靠区分正确性,常见缓解技术也无法解决该问题。

AI 中文摘要

大语言模型(LLMs)正越来越多地被用于自动化代码生成,但生成的程序可能在语法上看似合理,却仍无法通过基于执行的正确性检查。现有的验证方法,如测试和程序分析,仍然至关重要,但往往不完整、成本高昂,或仅在生成后应用。因此,模型衍生的不确定性自然成为一种早期可靠性信号。本文研究代码LLMs中的过度自信困境,即生成的错误程序往往具有与正确程序相当的 token 级置信度。我们在四个开源代码模型和三个基于执行的基准上研究这一困境。我们的分析首先调查现有的不确定性指标是否为代码生成中执行正确性提供可靠的代理。然后,我们在全局和局部 token 层面表征过度自信,探究错误程序在置信度和熵摘要(包括选择性生成和指令调优的限制)下是否仍与正确程序无法区分。最后,我们评估常见的缓解策略是否能减少这种失败模式。我们的研究得出四个发现:第一,现有的不确定性信号仅提供部分且依赖于模型的执行失败证据;第二,过度自信在程序和 token 层面均存在,基于不确定性的选择并不能始终提高接受集的准确率;第三,指令调优可增加失败生成的确定性,但无法始终提高正确性区分度;第四,常见的缓解技术可改善可靠性的特定方面,但无法可靠解决过度自信失败。我们的探索性潜在分析表明,隐藏表征可能编码输出置信度未暴露的与正确性相关的信号。

英文摘要

Large language models (LLMs) are increasingly used for automated code generation, but generated programs can appear syntactically plausible while still failing execution-based correctness checks. Existing validation methods, such as testing and program analysis, remain essential but are often incomplete, costly, or applied only after generation. Model-derived uncertainty is therefore a natural early reliability signal. This paper studies the dilemma of overconfidence in code LLMs where incorrect programs are often generated with token-level confidence comparable to correct programs. We study this dilemma across four open-source code models and three execution-based benchmarks. Our analysis begins by investigating whether existing uncertainty metrics provide reliable proxies for execution correctness in code generation. We then characterize overconfidence at both global and local token levels, asking whether incorrect programs remain indistinguishable from correct ones under confidence and entropy summaries, including selective generation and the limits of instruction tuning. Finally, we evaluate whether common mitigation strategies reduce this failure mode. Our study yields four findings. First, existing uncertainty signals provide only partial and model-dependent evidence of execution failure. Second, overconfidence persists at both program and token levels, and uncertainty-based selection does not consistently improve accepted-set accuracy. Third, instruction tuning can increase certainty on failing generations without consistently improving correctness discrimination. Fourth, common mitigation techniques improve specific aspects of reliability but do not reliably resolve overconfident failure. Our exploratory latent analysis suggests that hidden representations may encode correctness-related signals that output confidence does not expose.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑