arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23341cs.SE

DPIAgent:面向智能体复现测试用例生成的Divide、Protocol、Isolate方法

DPIAgent: Divide, Protocol, Isolate for Agentic Reproduction Test Generation

Hao Liu, Steven Liu, Xin Zhang, Jane Luo, Yu Kang, Jie Wu, Fangkai Yang, Yangyu Huang, Pengfei Gao, Scarlett Li, Yan Lu

首次发表
浏览论文内容

中文总结 AI 辅助

DPIAgent是基于Divide、Protocol、Isolate原则的结构化智能体框架,将复现测试用例生成拆分为单目标阶段,在SWT-Bench Verified上优于七个基线方法,GPT-5成功率达81.76%,添加测试选择后升至86.17%,通用性良好。

中文摘要 AI 辅助

复现测试用例生成是自动化软件工程中的关键步骤,其目标是生成能够捕获已报告缺陷的“先失败后通过”测试用例。现有的智能体方法将该任务视为一个整体循环,却忽略了该任务本质上包含性质截然不同的两个子任务:诊断根本原因和编写“先失败后通过”的测试用例。若不进行显式分离,智能体将面临目标不明确的复合目标,进而导致目标漂移。本文提出DPIAgent,这是一个基于Divide(拆分)、Protocol(协议)、Isolate(隔离)三项原则构建的结构化智能体框架,旨在缓解复合目标的模糊性和目标漂移问题:它将任务拆分为缺陷探索和测试用例生成两个单目标阶段;强制实施交接协议,记录诊断结果和测试计划,防止上下文丢失;并通过为每个阶段定制工具集来隔离其动作空间,避免无关工具误导执行。在SWT-Bench Verified基准测试中,DPIAgent在三种主干大语言模型(LLM)上的表现优于七个基线方法。仅使用DPI方法时,其在GPT-5上的成功率达到81.76%,是开源方法中报告的最高水平,在GPT-5-Mini上比最强基线方法高出11.88个百分点;添加测试选择机制后,成功率进一步提升至86.17%。分析表明,架构结构和主干能力是互补维度而非替代关系,证明了DPI在不同模型类别间的通用性。

英文摘要

Reproduction test generation, producing a failing-then-passing test that captures a reported bug, is a critical step in automated software engineering. Existing agentic methods treat this as a monolithic loop, despite the task inherently comprising two subtasks of distinct nature: diagnosing the root cause and writing a fail-to-pass test. Without explicit separation, the agent faces a compound objective with underspecified intermediate goals, leading to goal drift. We propose DPIAgent, a structured agentic framework built on three principles, Divide, Protocol, Isolate (DPI), that mitigates compound-objective ambiguity and goal drift: it Divides the task into single-objective phases of defect exploration and test generation; enforces a handoff Protocol that records the diagnosis and test plan, preventing context loss; and Isolates each phase's action space by tailoring the toolset to its task, preventing irrelevant tools from misleading execution. On SWT-Bench Verified, DPIAgent outperforms seven baselines across three backbone LLMs. With DPI alone it reaches 81.76% success rate on GPT-5, the highest reported among open-source methods, gaining up to 11.88 points over the strongest baseline on GPT-5-Mini; adding test selection further raises it to 86.17%. Our analysis shows that architectural structure and backbone capability are complementary axes rather than substitutes, demonstrating DPI's generalizability across model classes.

↑