发表机构
Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ExeCRE是基于执行一致性的代码可靠性估计框架,可减少自校正代码生成中的误导性反馈,提升有效性与稳定性,在GPT-5.2搭配LiveCodeBench及数学推理场景均有显著收益。
AI 中文摘要
大语言模型(LLMs)在代码生成领域已取得显著进展,但在需要复杂算法或实现的挑战性任务上仍存在不足。近期方法日益将代码执行作为反馈,尤其是在自校正流程中,这类流程会从生成的代码中构建验证信号。然而,这些流程常依赖可靠性未知的监督信号,可能引入误导性反馈、不必要的修正以及错误的最终答案。为解决该问题,我们提出ExeCRE——一个基于执行一致性引导的代码可靠性估计框架。ExeCRE并非通过测试或LLM反馈评判候选代码,而是通过对大量随机生成输入的执行输出进行统计分析,估计代码的可靠性;它收集生成输入的执行输出,将其投影为一致性信号,并应用Dawid-Skene模型推断潜在的代码可靠性。我们将ExeCRE集成到代码生成的自校正流程中,实验表明ExeCRE可持续提升有效性与稳定性,同时大幅减少误导性修正信号:在GPT-5.2搭配LiveCodeBench的设置下,已正确代码的平均误导性反馈案例数从代表性自校正基线的113.2降至ExeCRE的14.0。作为额外研究,我们将相同的可靠性估计策略应用于基于代码的数学推理,也观察到类似收益。这些结果表明ExeCRE可让基于执行的流程中更可靠地使用生成代码。
英文摘要
Large language models (LLMs) have made notable progress in code generation, but they still struggle on challenging tasks that require sophisticated algorithms or complex implementations. Recent methods increasingly use code execution as feedback, especially in self-correction pipelines that construct verification signals from generated code. However, these pipelines often depend on supervision signals whose reliability is unknown, which can introduce misleading feedback, unnecessary revisions, and incorrect final answers. To address this issue, we propose ExeCRE, an Execution-Consistency guided code Reliability Estimation framework. Instead of judging candidate code by tests or LLM feedback, ExeCRE estimates code reliability by statistically analyzing consistency patterns in execution outputs over a large number of randomly generated inputs. It collects execution outputs over generated inputs, projects them into consistency signals, and applies the Dawid-Skene model to infer latent code reliability. We integrate ExeCRE into self-correction for code generation. Experiments show that ExeCRE consistently improves both effectiveness and stability, while substantially reducing misleading correction signals. Under GPT-5.2 on LiveCodeBench, the average number of misleading feedback cases on already correct code drops from 113.2 with a representative self-correction baseline to 14.0 with ExeCRE. As an additional study, we apply the same reliability estimation strategy to code-based mathematical reasoning and observe similar benefits. These results suggest that ExeCRE enables more reliable use of generated code in execution-based pipelines.
Comments13 pages, 5 figures. Accepted at the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE 2026)