发表机构
Pi School(Pi学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RefactorPlatform 是一款开源工具,用于评估仓库级重构智能体,通过控制模型主干、执行机制等维度,在 RefactorBench 任务上验证了 AST 分块、检索增强等策略的效果,推动了评估的可复现性与可审计性。
AI 中文摘要
仓库级重构要求编码智能体在不改变程序行为的前提下,将单个变更传播到多个相互依赖的文件中,但据我们所知,现有工具均未分离出决定智能体在该任务上成功的设计选择。我们提出 RefactorPlatform,这是一个开源评估工具,它固定环境并显式调整每个设计维度:模型主干(通过 OpenRouter 和 GitHub Copilot CLI)、执行机制(基线、检索增强和多智能体)以及提示特异性。每次运行都在隔离工作区中执行,具备实时终端流、每任务的 token、差异和文本记录、基于 AST 的验证,以及可导出的遥测数据以用于审计和复现。我们在四个模型系列的 100 个多文件 RefactorBench 任务上展示该平台,说明其支持的分析:在所有提示模式下,感知 AST 的分块方法比朴素 token 窗口分块方法性能高 25-30%,而朴素检索则低于无检索基线;精简的检索增强单智能体(准确率 86%)在匹配任务上优于我们评估的子智能体配置(准确率 66%),且在检索下失败的任务在委派下也未通过;检索的准确率提升抵消了其 token 开销,使每次成功重构的成本保持不变。RefactorPlatform 已开源,旨在使重构智能体的评估可复现和可审计。
英文摘要
Repository-scale refactoring requires coding agents to propagate a single change across many interdependent files without altering program behavior, yet to our knowledge no existing harness isolates the design choices that determine agent success on this task. We present RefactorPlatform, an open-source evaluation harness that holds the environment fixed and varies each design axis explicitly: model backbone (via OpenRouter and GitHub Copilot CLI), execution regime (baseline, retrieval-augmented, and multi-agent), and prompt specificity. Each run executes in an isolated workspace with live terminal streaming, per-task logging of tokens, diffs, and transcripts, AST-based verification, and exportable telemetry for audit and reproduction. Demonstrating the platform on 100 multi-file RefactorBench tasks across four model families, we illustrate the analyses it supports: AST-aware chunking outperforms naive token-window chunking by 25-30% across prompt modes, whereas naive retrieval falls below the retrieval-free baseline; a lean retrieval-augmented single agent (86%) beats the sub-agent configuration we evaluated (66%) on matched tasks with no task passing under delegation that fails under retrieval; and retrieval's accuracy gains absorb its token overhead, leaving cost per successful refactoring unchanged. RefactorPlatform is open-sourced to make refactoring-agent evaluation reproducible and auditable.
CommentsAccepted at EMNLP 2026 System Demonstrations