AI 中文总结
研究针对大语言模型在自动程序修复中的局限,提出CT-Repair框架,通过代码属性图和时间执行图表示证据,利用三阶段过滤管道及多视角智能体分析错误并生成修复策略,实验证明该方法能有效提高修复有效性。
AI 中文摘要
大语言模型改进了自动程序修复,但存在两个局限性。原始执行跟踪通常太大且重复,无法作为有效的模型上下文。重复的补丁采样可能产生不同实现,但没有产生不同的根本原因假设或修复策略。我们提出了CT-Repair,一个智能程序修复框架,将静态和动态证据表示为可查询的代码属性图(CPG)和时间执行图(TEG)。CT-Repair应用三阶段过滤管道来构建紧凑的TEG。三个有限状态机引导的智能体从静态、动态和混合视角分析每个错误,并独立产生基于证据的修复策略。一个策略引导的生成过程将这些策略实例化为候选补丁,并使用验证反馈来完善最有希望的策略。我们在来自Defects4J v3.0的854个Java错误上评估了CT-Repair。在混合模型配置中,CT-Repair正确修复了489个错误。在受控的GPT-5.4-mini配置下,它修复了388个错误,分别比ReinFix和RepairAgent多19个和30个。三种证据视角的联合比最强的单个视角多修复了99个错误。过滤管道还压缩了运行时证据,执行过滤平均将候选方法范围缩小了94.85%,行为过滤进一步将保留的运行时记录减少了55.97%。这些结果表明,结构化运行时证据和多视角推理可以提高修复有效性,而不依赖于更大的补丁生成预算。
英文摘要
Large language models (LLMs) have improved automated program repair (APR), but two limitations remain. First, raw execution traces are often too large and repetitive to serve as effective model context. Second, repeated patch sampling may produce different implementations without yielding distinct root-cause hypotheses or repair strategies. We present CT-Repair, an agentic APR framework representing static and dynamic evidence as queryable Code Property Graph (CPG) and Temporal Execution Graph (TEG). CT-Repair applies a three-stage filtering pipeline to construct compact TEGs. Three finite-state-machine-guided agents analyze each bug from static, dynamic, and hybrid perspectives and independently produce evidence-grounded repair strategies. A strategy-guided generation procedure instantiates these strategies as candidate patches and uses validation feedback to refine the most promising strategy. We evaluate CT-Repair on 854 Java bugs from Defects4J v3.0. In the mixed-model configuration, CT-Repair correctly repairs 489 bugs. Under a controlled GPT-5.4-mini configuration, it repairs 388 bugs, 19 and 30 more than ReinFix and RepairAgent, respectively. The union of the three evidence perspectives repairs 99 more bugs than the strongest individual perspective. The filtering pipeline also compacts runtime evidence, with execution filtering narrowing the candidate method scope by 94.85% on average and behavior filtering further reducing retained runtime records by 55.97%. These results show that structured runtime evidence and multi-perspective reasoning can improve repair effectiveness without relying solely on a larger patch-generation budget.
Comments12 pages, 5 figures, 10 tables