发表机构
Navers Lab; Tsinghua University(Navers实验室; 清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出SWE Refactor Bench基准,通过三阶段评估发现前沿编码智能体仅5.4%能完成全仓库迁移,且迁移完整性与行为正确性是不同能力,该基准可作为可靠全仓库迁移编码智能体的测试平台。
AI 中文摘要
现代软件系统经过数十年开发会累积技术债务,这使得迁移成本高昂且大多依赖人工完成。随着编码智能体在漏洞修复方面的能力日益增强,它们能否自主完成这类迁移?现有基准无法回答该问题,因为它们仅评估行为正确性,而非迁移是否实际发生,这导致出现了一种简单的“取巧”方式:智能体复制原始实现以通过测试,我们将此称为“盲目性”。为解决该问题,我们推出SWE Refactor Bench,这是一个包含20项全仓库迁移任务的基准,涵盖4种技术债务类型。我们采用三阶段评估协议,同时衡量迁移完整性和行为正确性:(1)迁移审计(Migration Audit)验证迁移是否实际发生;(2)行为测试(Behavioural Tests)通过固定测试套件衡量正确性;(3)智能体验证(Agentic Verification)使用6个独立编码智能体,为隐藏的行为差异生成针对性测试。在8个前沿模型和26种模型-工作量配置的520次运行中,仅28次(5.4%)通过全部三个阶段,20项任务中有13项未获得可接受的解决方案,最优模型claude-opus-5的得分为47.0/100。迁移完整性和行为正确性是两种不同的能力:部分运行跳过迁移以保留行为,在迁移审计阶段被拦截;多数运行尝试迁移但破坏了行为,在行为测试阶段被拦截。智能体无法实现完美迁移:在通过迁移审计的340次运行中,58%达到99%的固定检查项,仅26%达到100%。智能体在不同迁移类别上的能力存在差异:在构建工具链重写任务上得分为31.4,而在语言重写任务上仅得5.6。综上,这些发现表明SWE Refactor Bench是开发用于可靠全仓库迁移的编码智能体的严格测试平台。
英文摘要
Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As coding agents become increasingly capable at bug fixing, can they autonomously perform such migrations? Existing benchmarks cannot answer this question because they evaluate only behavioural correctness, not whether the migration actually occurred. This leads an easy hack: agents copy the original implementation to make tests pass. We call this Blindness. To address this problem, we introduce SWE Refactor Bench, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt. A three-stage evaluation protocol measures both migration completeness and behavioural correctness. (1) Migration Audit verifies that the migration occurred. (2) Behavioural Tests measure correctness with a fixed test suite. (3) Agentic Verification uses 6 independent coding agents to generate targeted tests for hidden behavioural differences. Across 520 runs from 8 frontier models and 26 model-effort configurations, only 28 of 520 runs ($5.4\%$) pass all three stages, 13 of the 20 tasks receive no accepted solution, and the best model (claude-opus-5) scores $47.0/100$. Migration completeness and behavioural correctness are distinct abilities: a few runs preserve behaviour by skipping the migration and are stopped at Migration Audit; most attempt it and break behaviour, and are stopped at Behavioural Tests. Agents cannot deliver a perfect migration: among the 340 runs that pass Migration Audit, $58\%$ reach $99\%$ of the fixed checks, yet only $26\%$ reach $100\%$. Agent capability differs across migration categories: agents score $31.4$ on build toolchain rewrites but only $5.6$ on language rewrites. Together, these findings position SWE Refactor Bench as a rigorous testbed for developing coding agents for reliable whole-repository migrations.