arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于内省不确定性估计的LLM代码生成

Introspective Uncertainty Estimation for LLM-Based Code Generation

Thomas Klassert

arXiv 2609.13975首次发表:更新:

发表机构

Hochschule RheinMain(莱茵美因应用技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本论文研究基于LLM隐藏状态的内省不确定性估计,用于代码生成正确性评估,发现静态单令牌探针在响应级有效,支持两阶段风险筛查与故障定位流程。

AI 中文摘要

大型语言模型(LLMs)越来越多地用于代码生成,但可能产生流畅但功能不正确的输出,这限制了其在实际软件工程工作流程中的可信度。本论文研究基于LLM内部隐藏状态表示的内省不确定性估计(IUE)是否能在代码生成任务的响应级和行级可靠地指示正确性。目标是确定隐藏状态在多大程度上编码功能代码正确性的信息,以及如何利用这些信息进行实际风险评估和故障定位。在方法上,本论文将LiveCodeBench(LCB)和BigCodeBench(BCB)上的响应级评估与从错误程序派生令牌级和行级标签的增强流水线相结合。在此设置中,它比较了静态和动态响应级特征,评估了跨任务、编程领域和令牌位置的泛化能力,并研究了行级故障定位。结果表明,隐藏状态包含强烈的响应级正确性信号。静态单令牌探针表现最佳,而更复杂的动态策略没有产生一致的收益。虽然跨任务、领域和令牌位置的泛化是可行的,但针对真实世界软件项目的设置相关退化在很大程度上仍然存在。在细粒度上,行级预测比响应级估计困难得多。然而,在已知错误程序的条件定位设置中,Top-K故障点排序仍然有效。总体而言,研究结果表明,隐藏状态是估计功能代码正确性的稳健且信息丰富的资源,支持将响应级风险筛查与有针对性的行级优先级排序相结合的两阶段工作流程。

英文摘要

Large Language Models (LLMs) are increasingly used for code generation but can produce fluent yet functionally incorrect outputs, which limits trust in their usage for practical software engineering workflows. This thesis investigates whether Introspective Uncertainty Estimation (IUE), based on internal hidden-state representations of LLMs, can reliably indicate correctness at the response and line levels for code generation tasks. The objective is to determine the extent to which hidden states encode information about functional code correctness and how this can be leveraged for practical risk assessment and fault localization. Methodologically, this thesis combines response-level evaluation on LiveCodeBench (LCB) and BigCodeBench (BCB) with an augmentation pipeline that derives token- and line-level labels from incorrect programs. In this setup, it compares static and dynamic response-level features, evaluates generalization across tasks, programming domains, and token positions, and studies line-level fault localization. The results show that hidden states contain a strong response-level correctness signal. Static single-token probes perform best, while more elaborate dynamic strategies yield no consistent gains. While generalization across tasks, domains, and token positions is feasible, setting-dependent degradation largely remains for real-world software projects. At a fine granularity, line-level prediction is substantially harder than response-level estimation. However, in a conditional localization setup with known-incorrect programs, Top-K point-of-failure ranking remains effective. Overall, the findings suggest that hidden states are a robust and informative resource for estimating functional code correctness, supporting a two-stage workflow that combines response-level risk screening with targeted line-level prioritization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑