arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReProAgent:从问题报告中通过工具增强的多阶段代理生成错误重现测试

ReProAgent: Tool-Augmented Multi-Stage Agentic Generation of Bug Reproduction Tests from Issue Reports

Quanjun Zhang, Yi Zheng, Ye Shang, Weifeng Sun, Haichuan Hu, Chunrong Fang, Zhenyu Chen, Liang Xiao

arXiv 2607.09123首次发表:更新:

AI 中文总结

研究提出ReProAgent多阶段代理框架,从问题报告生成错误重现测试,将任务分解为四个阶段,集成多种工具支持各阶段,实验表明其能成功重现较多问题,优于基线,还可跨模型泛化并提升下游问题解决性能。

AI 中文摘要

重现测试有助于开发人员确认报告的问题并为问题解决提供可执行反馈,但开源项目中的问题报告很少包含此类测试。近期研究探索用大语言模型从问题报告生成测试,但现有方法大多依赖基于提示的管道。本文提出ReProAgent,一个用于从问题报告生成重现测试的多阶段代理框架。它将任务分解为四个代理阶段:错误定位、根本原因分析、测试计划和测试生成。通过集成特定任务工具、文本和仓库图的上下文检索以及与执行环境的运行时交互来支持这些阶段。实验表明ReProAgent分别成功重现了58.43%和70.30%的问题,优于所有基线,每个实例平均成本为0.14美元,还能跨多种骨干语言模型进行泛化并提高下游问题解决性能。

英文摘要

Reproduction tests help developers confirm reported issues and provide executable feedback for issue resolution, yet issue reports in open-source projects rarely include such tests. Recent studies have explored generating issue reproduction tests from issue reports with large language models, but existing approaches largely rely on prompt-based pipelines that retrieve textual context and generate tests. This limits their ability to understand how reported issues behave in repository-scale codebases and to flexibly organize the construction of reproduction tests. In this paper, we propose ReProAgent, a multi-stage agent framework for reproduction test generation from issue reports. ReProAgent decomposes the task into four agent stages: bug localization, root cause analysis, test planning, and test generation. To support these stages, ReProAgent integrates task-specific tools for task decomposition and reflection, context retrieval from both textual sources and repository graphs, and runtime interaction with the execution environment. Experiments on SWT-bench-lite and SWT-bench-verified show that ReProAgent successfully reproduces 58.43% and 70.30% of issues, outperforming all baselines, with an average cost of $0.14 per instance. For example, when equipped with GPT-5-mini, ReProAgent exceeds OpenHands with the same backbone by 20.43 and 7.90 percentage points, respectively. ReProAgent also generalizes across multiple backbone LLMs and improves downstream issue resolution performance when integrated with existing repair approaches.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑