发表机构
Massachusetts Institute of Technology; Rensselaer Polytechnic Institute; IBM; Michigan State University(麻省理工学院; 罗切斯特理工学院; IBM公司; 密歇根州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对仓库级代码优化难题,提出PerfAgent分析器引导迭代优化方法,该方法能让编码代理找到热点并改进,在两个优化基准测试中大幅提升专家匹配补丁率,且以低成本超越基线。
AI 中文摘要
大型语言模型(LLM)代理在面向正确性的仓库级任务上表现良好,但在仓库级代码优化方面仍有困难。当前代理常错过隐藏瓶颈,加速有限或未充分测试补丁。我们提出PerfAgent,一种分析器引导、验证器参与的工作流程,能让现成编码代理找到真正热点,超越首次通过补丁进行改进,并利用分析器证据决定下一步优化。在两个优化基准测试中,PerfAgent比OpenHands with GPT-5.1的专家匹配补丁率提高一倍多,还以更低成本超越了神谕五选最佳基线。
英文摘要
Large language model (LLM) agents now perform well on correctness-oriented repository-level tasks, including SWE-Bench issue resolution and feature implementation in real codebases. However, they still struggle with repository-level code optimization, which requires preserving behavior while improving runtime performance. Passing tests is not enough in this setting; a patch must preserve behavior, implement code optimization, and approach expert speedups. Current agents often miss bottlenecks hidden behind abstraction layers and native extensions, stop after shallow speedups, or insufficiently test the code patches that thus may silently break edge cases. We present PerfAgent, a profiler-guided, verifier-in-the-loop workflow that gives an off-the-shelf coding agent the feedback needed to find real hotspots, improve beyond the first passing patch, and use profiler evidence rather than timing alone to decide what to optimize next. On two challenging optimization benchmarks, GSO and SWE-fficiency-Lite, PerfAgent more than doubles the rate of expert-matching patches over OpenHands with GPT-5.1, improving from 19.6% to 39.2% on GSO and from 26% to 74% on SWE-fficiency-Lite. It also surpasses an oracle best-of-five baseline at substantially lower cost, showing that the gains come from better feedback rather than additional test-time sampling.