基于大语言模型的程序推理何时正确?基于大语言模型的代码推理的补全语义
When is LLM-Based Program Reasoning Correct? A Completion Semantics for LLM-Based Code Inference
浏览论文内容
中文总结 AI 辅助
研究基于大语言模型的程序推理何时正确,引入补全语义,将不完整程序形式化,定义存在性推理正确性。以见证生成工作流程实例化方法,在真实任务上评估,能区分合理与不合理推理,为验证不完整程序推理提供实用机制。
中文摘要 AI 辅助
由于令牌和认知限制,大语言模型(LLMs)通常对不完整的代码片段/提示进行程序推理,而非完整程序。这种推理依赖于对省略代码和上下文的假设,因此程序片段推理的意义并非绝对,而是取决于描述片段如何细化为完整程序的隐式补全模型。本文引入基于LLM的程序推理的补全语义,将不完整程序形式化为表示可能细化的空间,并定义相对于补全模型的存在性推理的正确性。基于此,当模型中存在见证错误的补全时,报告的错误就是正确的。这种观点解释了为何许多LLM生成的报告既非简单正确也非错误,而是取决于对省略上下文的假设。我们以见证生成工作流程的形式实例化了方法,通过构建原始程序片段的可执行细化来具体化推理背后的补全。见证既作为存在性断言的证据,也作为揭示支持这些断言所需假设的机制。我们在真实世界中由LLM生成的错误报告和程序分析任务上评估了方法。结果表明,见证生成有效地将由合理补全支持的推理与需要不切实际假设的推理区分开来,为验证对不完整程序的推理提供了实用机制。
英文摘要
Due to token and cognitive limits, Large Language Models (LLMs) typically perform program reasoning over incomplete code fragments/prompts rather than complete programs. Such reasoning therefore must rely on {assumptions about omitted code and context. As a result, the meaning of an inference over a program fragment is not absolute, but depends on an implicit completion model describing how the fragment may be refined into a complete program. In this paper, we introduce completion semantics for LLM-based program reasoning. We formalize incomplete programs as denoting a space of possible refinements and define the correctness of existential inferences relative to a completion model. Under this view, a reported bug is correct whenever there exists a completion within the model that witnesses the bug. This perspective explains why many LLM-generated reports are neither simply correct nor incorrect, but instead depend on assumptions about omitted context. We have instantiated our approach in the form of a witness-generation workflow that concretizes completions underlying an inference by constructing executable refinements of the original program fragment. Witnesses serve both as evidence for existential claims and as a mechanism for exposing the assumptions required to support them. We evaluate our approach on real-world LLM-generated bug reports and program-analysis tasks. Our results show that witness generation effectively distinguishes inferences supported by plausible completions from those requiring unrealistic assumptions, providing a practical mechanism for validating reasoning over incomplete programs.
发表机构
- Nanjing University(南京大学)
- National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。