DPIAgent:面向智能体复现测试用例生成的Divide、Protocol、Isolate方法
DPIAgent: Divide, Protocol, Isolate for Agentic Reproduction Test Generation
浏览论文内容
中文总结 AI 辅助
DPIAgent是基于Divide、Protocol、Isolate原则的结构化智能体框架,将复现测试用例生成拆分为单目标阶段,在SWT-Bench Verified上优于七个基线方法,GPT-5成功率达81.76%,添加测试选择后升至86.17%,通用性良好。
中文摘要 AI 辅助
复现测试用例生成是自动化软件工程中的关键步骤,其目标是生成能够捕获已报告缺陷的“先失败后通过”测试用例。现有的智能体方法将该任务视为一个整体循环,却忽略了该任务本质上包含性质截然不同的两个子任务:诊断根本原因和编写“先失败后通过”的测试用例。若不进行显式分离,智能体将面临目标不明确的复合目标,进而导致目标漂移。本文提出DPIAgent,这是一个基于Divide(拆分)、Protocol(协议)、Isolate(隔离)三项原则构建的结构化智能体框架,旨在缓解复合目标的模糊性和目标漂移问题:它将任务拆分为缺陷探索和测试用例生成两个单目标阶段;强制实施交接协议,记录诊断结果和测试计划,防止上下文丢失;并通过为每个阶段定制工具集来隔离其动作空间,避免无关工具误导执行。在SWT-Bench Verified基准测试中,DPIAgent在三种主干大语言模型(LLM)上的表现优于七个基线方法。仅使用DPI方法时,其在GPT-5上的成功率达到81.76%,是开源方法中报告的最高水平,在GPT-5-Mini上比最强基线方法高出11.88个百分点;添加测试选择机制后,成功率进一步提升至86.17%。分析表明,架构结构和主干能力是互补维度而非替代关系,证明了DPI在不同模型类别间的通用性。
英文摘要
Reproduction test generation, producing a failing-then-passing test that captures a reported bug, is a critical step in automated software engineering. Existing agentic methods treat this as a monolithic loop, despite the task inherently comprising two subtasks of distinct nature: diagnosing the root cause and writing a fail-to-pass test. Without explicit separation, the agent faces a compound objective with underspecified intermediate goals, leading to goal drift. We propose DPIAgent, a structured agentic framework built on three principles, Divide, Protocol, Isolate (DPI), that mitigates compound-objective ambiguity and goal drift: it Divides the task into single-objective phases of defect exploration and test generation; enforces a handoff Protocol that records the diagnosis and test plan, preventing context loss; and Isolates each phase's action space by tailoring the toolset to its task, preventing irrelevant tools from misleading execution. On SWT-Bench Verified, DPIAgent outperforms seven baselines across three backbone LLMs. With DPI alone it reaches 81.76% success rate on GPT-5, the highest reported among open-source methods, gaining up to 11.88 points over the strongest baseline on GPT-5-Mini; adding test selection further raises it to 86.17%. Our analysis shows that architectural structure and backbone capability are complementary axes rather than substitutes, demonstrating DPI's generalizability across model classes.