真实仓库中可信的运行时错误修复:基准与护栏
Trustworthy Runtime Error Healing in Real-World Repositories: A Benchmark and Guardrail
浏览论文内容
中文总结 AI 辅助
本文提出HealBench基准和HealGuard护栏,用于在真实仓库中安全地进行LLM运行时错误修复,实验显示现有智能体能修复部分崩溃,但HealGuard以较高误报率检测潜在风险。
中文摘要 AI 辅助
运行时错误修复通过生成代码来修复崩溃程序的实时运行状态,从而使程序能够继续运行。近期研究表明,大型语言模型(LLM)能够生成此类修复代码,但相关评估仅在小规模竞赛程序上进行,并且在实时进程中执行LLM生成的代码会引发尚未解决的安全问题。本文致力于将基于LLM的运行时错误修复推向真实世界仓库的实际应用。首先,我们构建了HealBench基准,包含来自18个真实仓库的265个运行时错误,每个错误均配有修复版本上的参考执行。HealBench还提供了一个统一框架,使LLM智能体能够利用跨文件上下文和实时运行状态进行修复。随后,我们设计了HealGuard,要求修复代码使用HealCore(Python的一个可分析子集)编写,并利用静态和动态污点分析来检查修复所改变的状态是否触及开发者保护的操作。我们评估了一种专用修复方法和三种通用编码智能体,并搭配三种骨干LLM。最佳设置在38.11%的实例中恢复执行,并在28.68%的实例中通过目标测试,表明现有智能体已经能够修复相当比例的真实仓库级崩溃。然而,在通过的执行中,HealGuard标记出17.4%的实例,其修复改变的状态可能触及受保护操作。在684个受控案例中,HealGuard检测出所有不安全案例,但代价是68.42%的误报率。
英文摘要
Runtime error healing lets a crashed program continue by generating code that repairs its live runtime state. Recent work shows that LLMs can generate such healing code, but it is evaluated only on small competition programs, and executing LLM-generated code inside a live process raises safety concerns that remain unaddressed. In this paper, we take LLM-based runtime healing toward practical use in real-world repositories. We first build HealBench, a benchmark of 265 runtime errors from 18 real-world repositories, each paired with a reference execution on the patched version. HealBench also provides a unified framework that lets LLM agents heal with cross-file context and live runtime state. We then design HealGuard, which requires healing code to be written in HealCore, an analyzable subset of Python, and uses static and dynamic taint analysis to check whether state changed by healing reaches operations protected by developers. We evaluate a dedicated healing method and three general coding agents with three backbone LLMs. The best setting resumes execution in 38.11% of instances and passes the target test in 28.68%, showing that existing agents can already heal a meaningful share of real repository-level crashes. However, among executions that pass, HealGuard flags 17.4% whose healing-changed state may reach a protected operation. On 684 controlled cases, HealGuard detects all unsafe cases, at the cost of a 68.42% false positive rate.
发表机构
- Sun Yat-sen University(中山大学)
- Singapore Management University(新加坡管理大学)
- Monash University(莫纳什大学)
机构由 AI 辅助整理,请以论文原文为准。