arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10242cs.SEcs.AI

TestGRAD:通过失败模式动量演化测试套件以用于SWE-Agent集成

TestGRAD: Evolving Test Suites via Failure Pattern Momentum for SWE-Agent Ensemble

  • University of Manitoba(曼尼托巴大学)
  • Huawei Canada(华为加拿大)
  • Huawei Technologies(华为技术)

机构由 AI 辅助整理,请以论文原文为准。

Pengfei He, Jiayuan Zhou, Shaowei Wang, Ruiqi Pan

中文总结 AI 辅助

TestGRAD提出一种基于失败模式动量的测试套件演化框架,通过差分损失和完整CRUD梯度优化测试生成,在SWE-bench Verified上以4智能体集成实现84.2%的Pass@1,超越基线3.6个百分点。

中文摘要 AI 辅助

SWE-agent集成通过结合来自具有互补优势的不同智能体的候选补丁来提高问题解决能力。因此,核心问题是基于测试的选择:生成测试、执行候选补丁并识别最佳补丁。我们将此过程形式化为测试空间优化:演化一个可执行的仓库测试套件,直到它能够区分竞争性补丁。现有的测试生成方法是有限的优化器。它们通常缺乏用于集成选择的显式损失,通过不完整的方向进行优化(这些方向大多创建新测试或删除旧测试),并且执行一次性生成,没有来自重复失败的反馈。受带动量的梯度下降启发,我们引入了TestGRAD,一个用于自动测试优化的框架。TestGRAD围绕三个概念展开。差分损失为优化器提供了明确的执行定义目标:有用的测试应通过行为区分候选补丁。完整CRUD梯度将更新方向从仅仅创建或删除测试扩展到读取现有测试基础设施、创建新测试、更新过时的断言,以及仅删除过时的测试。失败模式动量从记忆中挖掘频繁的失败序列,使优化器能够避免重复的非区分性方向,同时压缩失败历史上下文。在SWE-bench Verified上,TestGRAD在4智能体集成下达到84.2%的Pass@1,比最强基线(80.6%)绝对提高了3.6个百分点,同时将失败历史上下文压缩了超过100倍。

英文摘要

SWE-agent ensembles improve issue resolution by combining candidate patches from different agents with complementary strengths. The central problem is therefore test-based selection: generate tests, execute candidate patches, and identify the best patch. We formulate this process as test-space optimization: evolving an executable repository test suite until it distinguishes competing patches. Existing test-generation methods are limited optimizers. They usually lack an explicit loss for ensemble selection, optimize through incomplete directions that mostly create new tests or delete old ones, and perform one-off generation without feedback from repeated failures. Inspired by gradient descent with momentum, we introduce TestGRAD, a framework for automatic test optimization. TestGRAD centers on three concepts. Differential loss gives the optimizer an explicit execution-defined target: useful tests should separate candidate patches by behavior. Full CRUD gradients expand the update direction from merely creating or deleting tests to reading existing test infrastructure, creating new tests, updating stale assertions, and deleting only obsolete tests. Failure Pattern Momentum mines frequent failure sequences from memory, allowing the optimizer to avoid repeated non-discriminative directions while compressing the failure-history context. On SWE-bench Verified, TestGRAD achieves 84.2% Pass@1 with a 4-agent ensemble, outperforming the strongest baseline (80.6%) by an absolute improvement of 3.6 percentage points, while compressing failure-history context by over $100\times$.

补充信息

↑