arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Active-SWE:针对无问题报告的主动漏洞修复任务的编码智能体基准测试

Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports

Haobin Li, Ping Deng, Weizhong Qian, Liang Jiang, Zhenyu Huang, Mouxing Yang, Xi Peng

arXiv 2608.04682首次发表:更新:

发表机构

Sichuan University; University of Electronic Science and Technology of China(四川大学; 电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Active-SWE是首个针对无问题报告的主动多漏洞修复任务的编码智能体基准,实验显示现有顶尖编码智能体在该任务上性能有限。

AI 中文摘要

由大语言模型(LLM)驱动的编码智能体正越来越多地被应用于软件工程(SWE)场景,能够修复大型代码库中的特定漏洞。然而,现有的SWE基准通常假设始终存在带有详细信息的高质量问题报告,但在实践中,由于报告获取和整理的复杂性,这一假设极易不成立。为解决该问题,我们推出Active-SWE,这是一个用于评估编码智能体在无报告指导下主动发现并修复多个漏洞的基准,涵盖6个漏洞类别和8种语言的1663个任务。除了将研究重点从现有被动漏洞修复转向主动漏洞修复外,Active-SWE还通过将范围从修复特定记录漏洞扩展到多漏洞修复及潜在漏洞发现场景,实现更深入的评估。为构建Active-SWE,我们提出一种新颖的难度感知任务制定流程及双轨评估框架,以全面评估主动漏洞修复能力。大量实验表明,大多数最先进的编码智能体在主动漏洞修复任务中表现不佳,在定位和解决记录漏洞、处理多漏洞修复场景以及发现有效潜在漏洞方面性能有限。

英文摘要

Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a specific bug in large-scale codebase. However, existing SWE benchmarks typically assume that high-quality issue reports with detailed information are always available, which is easily violated in practice due to the complexity of report acquisition and curation. To address this, we introduce Active-SWE, a benchmark for evaluating coding agents on proactively discovering and fixing multiple bugs without report guidance, covering 1,663 tasks across six bug categories and eight languages. Beyond shifting the focus from existing reactive bug fixing to proactive bug fixing, Active-SWE enables a more in-depth evaluation by expanding the scope from fixing a specific recorded bug to multiple-bug fixing and potential bug discovery scenarios. To construct Active-SWE, we propose a novel difficulty-aware task formulation pipeline with a dual-track evaluation framework, facilitating comprehensive evaluation of proactive bug-fixing capability. Extensive experiments reveal that most state-of-the-art coding agents struggle with proactive bug-fixing tasks, demonstrating limited performance in locating and resolving recorded bugs, handling multiple bug fixing scenarios, and discovering valid potential bugs.

Comments24 pages, 17 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑