arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10613cs.SE

CausalRepair:通过双重切片弥合基于大语言模型的自动化程序修复中的因果差距

CausalRepair: Bridging the Causality Gap in Large Language Model-Based Automated Program Repair via Dual-Slicing

Linhao Wu, Yizhou Chen, Zhen Yang, Pengyu Xue, Dan Hao

首次发表
浏览论文内容

中文总结 AI 辅助

CausalRepair是基于最小因果上下文的对话驱动型自动化程序修复框架,通过双重切片策略构建因果相关上下文,在Defects4J数据集上修复313个漏洞,性能优于现有方法且修复成本更低。

中文摘要 AI 辅助

自动化程序修复(APR)近期从大语言模型(LLM)中获益,但其有效性高度依赖修复上下文。现有基于LLM的APR方法存在因果差距问题:测试上下文可能存在噪声或不完整,而通过静态分析得到的源上下文常包含不相关且未执行的代码,误导LLM无法识别真正的根本原因。为解决该问题,我们提出CausalRepair,这是一个基于最小因果上下文(即解释故障所需的必要依赖)的对话驱动型APR框架。CausalRepair采用双重切片策略:上下文感知的静态切片净化测试语义,基于执行轨迹的动态切片捕获源代码中精确的运行时依赖,二者共同构建紧凑、因果相关的上下文以指导迭代修复。我们使用DeepSeek-V3在Defects4J V1.2、V2.0及Defects4J-Trans上评估CausalRepair,其在Defects4J上成功修复313个漏洞,性能优于ReinFix、TSAPR等最先进方法,同时将平均修复成本降至每个漏洞0.029美元。

英文摘要

Automated Program Repair (APR) has recently benefited from Large Language Models (LLMs), yet their effectiveness heavily depends on repair context. Existing LLM-based APR methods suffer from a causality gap: test contexts can be noisy or incomplete, while source contexts derived from static analysis often contain irrelevant and unexecuted code, misleading LLMs from identifying the true root cause. To address this issue, we propose CausalRepair, a conversation-driven APR framework based on minimal causal context, i.e., the essential dependencies required to explain a failure. CausalRepair employs a dual-slicing strategy: context-aware static slicing purifies test semantics, while execution-trace-based dynamic slicing captures precise runtime dependencies in source code. Together, they construct compact, causally relevant contexts to guide iterative repair. We evaluate CausalRepair on Defects4J V1.2, V2.0, and Defects4J-Trans using DeepSeek-V3. CausalRepair correctly fixes 313 bugs on Defects4J, outperforming state-of-the-art approaches such as ReinFix and TSAPR, while reducing the average repair cost to $0.029 per bug.

↑