arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14863cs.SEcs.AIcs.DC

评估分布式系统中智能体的代码修复能力

Evaluating Agentic Code Repair Capabilities in Distributed Systems

  • University of Southern California(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

Yibo Yan, Huijuan Wang, Junzhou He, Yizhuo Liang, Shaoyu Wang, Huanchen Sun, Seo Jin Park

AI总结:

该研究推出分布式系统代码修复基准DDBench,评估10个LLM的代码修复能力,发现调试上下文可提升准确率且对强弱模型影响不同,分布式调试能凸显单进程基准未体现的推理能力。

AI中文摘要:

基于大语言模型(LLM)的编码智能体在单进程软件维护(SWE)任务上发展迅速,前沿模型在SWE-bench Verified上的准确率已集中在70%以上。然而,分布式系统调试仍是一个研究不足的领域:错误跨越进程、节点和协议交互,仅从源代码中难以恢复根本原因,且在非确定性交织情况下暴力探索不可行。这在LLM和智能体评估中留下两个空白:没有针对分布式系统错误的代码修复基准,也没有对照研究分离外部提供的调试上下文对智能体在这类任务上成功的影响程度。我们推出DDBench,一个从13个开源分布式系统中挖掘的60个历史错误组成的代码修复基准,分为三个难度等级。DDBench在两种匹配条件下评估每个案例:仅症状条件下,智能体仅收到错误症状和代码库;上下文增强条件下,智能体额外收到有限的调试上下文(日志、跟踪、运行时状态和针对性代码调查注释),从而将调试上下文的影响与模型能力分离。对10个LLM在DDBench上的评估揭示了几项发现:第一,分布式调试展现了单进程基准未凸显的推理维度:模型在DDBench最难案例集上的准确率跨度达61个百分点,15对顶级模型中有9对通过成对自举检验在p<0.05水平上存在显著差异;第二,有限调试上下文使总准确率提升18.1个百分点,且提升不对称:较弱模型提升准确率,较强模型提升效率;第三,调试上下文需要精心整理,因为即使是准确的调试上下文有时也会误导LLM。

英文摘要:

LLM-based coding agents have advanced rapidly on single-process SWE tasks, with frontier models now clustering in the high-70s on SWE-bench Verified. Distributed-system debugging, however, remains an under-explored regime: bugs span processes, nodes, and protocol interactions, with root causes rarely recoverable from source alone and brute-force exploration intractable across non-deterministic interleavings. This leaves two gaps in LLM and agent evaluation: no code-repair benchmark targets distributed-system bugs, and no controlled study isolates how much externally provided debugging context changes agent success on them. We introduce DDBench, a code-repair benchmark of 60 historical bugs mined from 13 open-source distributed systems, partitioned into three difficulty tiers. DDBench evaluates every case under two matched conditions: a symptom-only condition where the agent receives only the bug symptom and repository, and a context-augmented condition where it additionally receives a bounded debugging context (logs, traces, runtime state, and targeted code-investigation notes), isolating the effect of debugging context from model capability. The evaluation of ten LLMs on DDBench reveals several findings. First, distributed debugging exercises a reasoning dimension that single-process benchmarks do not surface: models' pass rates span 61 pp, and pairwise bootstrap separates 9 of 15 top-tier model pairs at p < 0.05 on DDBench's hardest case-set. Second, bounded debugging context lifts aggregate pass rate by +18.1 pp, and the lift is asymmetric: weaker models gain pass rate, while stronger models gain efficiency. Third, debugging context requires careful curation, as even faithful debugging context can sometimes mislead LLMs.

补充信息

↑