arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.09553cs.SE

为软件修复代理编写错误报告:哪些信息最重要?

Writing Bug Reports for Software Repair Agents: What Information Matters Most?

Vincenzo Luigi Bruno, Alessandro Giagnorio, Daniele Bifolco, Leon Wienges, Massimiliano Di Penta, Gabriele Bavota

AI总结:

研究软件开发中为软件修复代理编写错误报告的问题,通过对错误报告分类标注,用三个LLM主干运行mini-swe-agent,拟合回归模型,发现定位线索和建议修复等信息对代理成功更重要,传统对人类有用的信息作用较小。

AI中文摘要:

软件开发正日益朝着以代理为先的工作流程发展,错误报告成为代理的主要任务规范。然而,对于如何编写此类报告以最大化代理生成正确修复的机会知之甚少。我们从SWE-bench验证基准开始研究,手动对每个问题按变更类型分类,并标注句子的信息类型。聚焦441个错误报告问题,使用三个LLM主干运行mini-swe-agent。通过拟合二项式回归模型,控制混杂因素,估计每种信息类型与代理成功之间的增量关联。结果表明,定位线索与成功修复正相关,建议修复与通过概率有很强正相关。消融研究证实代理从传统上对人类有用的信息中受益较少,从暴露修复方向的句子中受益更多。

英文摘要:

Software development is increasingly moving toward agentic-first workflows. This includes AI agents responsible for generating initial fixes for submitted issue reports. In this setting, issue reports are no longer merely documentation for human maintainers; they become the primary task specification for the agent. However, little is known about how such reports should be written to maximize the agent's chances of producing a correct fix. We study what makes a bug report agent-ready. Starting from the SWE-bench Verified benchmark (i.e., a collection of 500 real repository issues with human-written gold patches and test suites for evaluating generated fixes) we manually classify each issue by change type (e.g., bug fix vs refactoring) and annotate each sentence with its information type, such as observed behavior, expected behavior, reproduction steps, localization cues, and suggested fixes. We focus on the 441 issues representing bug reports, and we run on them mini-swe-agent using three LLM backbones (i.e., GPT-5-mini, MiniMax M2.5, and Gemini 3 Flash). We then fit a binomial regression model to estimate the incremental association between each information type and agent success, controlling for confounding factors. Our results suggest that agentic-first reports benefit most from information that narrows the agent's search and repair space. Localization cues, such as references to affected code areas, are positively associated with successful repairs, while suggested fixes, expressed either in code or natural language, show some of the strongest positive associations with pass probability. An ablation study removing selected information types confirms that agents benefit less from information traditionally useful to humans, such as reproduction steps, and more from sentences that expose a repair direction, either through bug localization or a suggested fix.

↑