arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34359cs.AI

通过运行时程序状态推理改进代码大语言模型

Improving Large Language Models for Code through Runtime Program-State Reasoning

  • University of California, Santa Barbara(加州大学圣塔芭芭拉分校)
  • Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

Hongwei Li, Spandan Garg, Yufan Huang

AI总结:

本文提出通过两种运行时程序状态推理任务(有缺陷输入输出推理和前置条件后置条件推理)训练代码大语言模型,开发出Comet-9B,在多个基准上显著提升补丁生成、回归测试生成等能力。

AI中文摘要:

大语言模型在推理运行时程序状态方面接受的显式训练有限。我们研究训练模型推理运行时程序状态是否能提升下游软件工程能力。我们引入了两个互补的程序状态推理任务。有缺陷的输入-输出推理要求模型生成一个具体的输入,该输入能暴露有缺陷程序与隐藏的正确实现之间的行为差异,并预测由此产生的执行行为。前置条件-后置条件推理要求智能体以符号方式刻画触发缺陷的前置条件,预测预期的后置条件,解释其因果联系,并将这一推理实例化为可执行的回归测试。通过将这两个任务纳入分阶段的后期训练流程,我们开发了Comet-9B,一个基于Qwen3.5-9B Base的90亿参数语言模型。我们在仓库级补丁生成、回归测试生成和安全概念验证生成上评估了所得的检查点。在问题解决的有监督微调(SFT)基础上添加这两个程序状态推理任务,在SWE-bench Pro上成功率提高了7.25个百分点,在SWT-Bench Verified上提高了9.70个百分点。对这两个任务进行顺序强化学习,在SWE-bench Pro、SWT-Bench Verified和CyberGym上分别进一步提升了7.25、26.79和4.67个百分点。尽管仅有90亿参数,Comet-9B在SWE-bench Pro上取得了与所报告的GPT-5.2结果相当的分数,并在SWT-Bench Verified上匹配了基于GPT-4o的智能体所报告的成功率。

英文摘要:

Large language models receive limited explicit training in reasoning about runtime program states. We study whether training models to reason about runtime program states improves downstream software-engineering capabilities. We introduce two complementary program-state reasoning tasks. Buggy input-output reasoning requires a model to generate a concrete input that exposes a behavioral difference between a buggy program and a hidden correct implementation and to predict the resulting execution behavior. Precondition-postcondition reasoning requires an agent to symbolically characterize a bug-triggering precondition, predict the expected postcondition, explain their causal connection, and instantiate this reasoning as an executable regression test. By incorporating these two tasks into a staged post-training pipeline, we develop Comet-9B, a 9B language model based on Qwen3.5-9B Base. We evaluate the resulting checkpoints on repository-level patch generation, regression-test generation, and security PoC generation. Adding both program-state reasoning tasks to supervised fine-tuning (SFT) on issue resolution improves success rates by 7.25 percentage points on SWE-bench Pro and 9.70 points on SWT-Bench Verified. Sequential reinforcement learning on the two tasks yields further gains of 7.25, 26.79, and 4.67 percentage points on SWE-bench Pro, SWT-Bench Verified, and CyberGym, respectively. Despite having only 9B parameters, Comet-9B achieves a score comparable to the reported GPT-5.2 result on SWE-bench Pro and matches the reported success rate of a GPT-4o-based agent on SWT-Bench Verified.

↑