arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PerfAgent:用于仓库级代码优化的分析器引导迭代优化

PerfAgent: Profiler-Guided Iterative Refinement for Repository-Level Code Optimization

Ryan Deng, Yuanzhe Liu, Bastian Lipka, Yao Ma, Xuhao Chen, Tim Kaler, Jatin Ganhotra

arXiv 2607.19653首次发表:更新:

发表机构

Massachusetts Institute of Technology; Rensselaer Polytechnic Institute; IBM; Michigan State University(麻省理工学院; 罗切斯特理工学院; IBM公司; 密歇根州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对仓库级代码优化难题,提出PerfAgent分析器引导迭代优化方法,该方法能让编码代理找到热点并改进,在两个优化基准测试中大幅提升专家匹配补丁率,且以低成本超越基线。

AI 中文摘要

大型语言模型(LLM)代理在面向正确性的仓库级任务上表现良好,但在仓库级代码优化方面仍有困难。当前代理常错过隐藏瓶颈,加速有限或未充分测试补丁。我们提出PerfAgent,一种分析器引导、验证器参与的工作流程,能让现成编码代理找到真正热点,超越首次通过补丁进行改进,并利用分析器证据决定下一步优化。在两个优化基准测试中,PerfAgent比OpenHands with GPT-5.1的专家匹配补丁率提高一倍多,还以更低成本超越了神谕五选最佳基线。

英文摘要

Large language model (LLM) agents now perform well on correctness-oriented repository-level tasks, including SWE-Bench issue resolution and feature implementation in real codebases. However, they still struggle with repository-level code optimization, which requires preserving behavior while improving runtime performance. Passing tests is not enough in this setting; a patch must preserve behavior, implement code optimization, and approach expert speedups. Current agents often miss bottlenecks hidden behind abstraction layers and native extensions, stop after shallow speedups, or insufficiently test the code patches that thus may silently break edge cases. We present PerfAgent, a profiler-guided, verifier-in-the-loop workflow that gives an off-the-shelf coding agent the feedback needed to find real hotspots, improve beyond the first passing patch, and use profiler evidence rather than timing alone to decide what to optimize next. On two challenging optimization benchmarks, GSO and SWE-fficiency-Lite, PerfAgent more than doubles the rate of expert-matching patches over OpenHands with GPT-5.1, improving from 19.6% to 39.2% on GSO and from 26% to 74% on SWE-fficiency-Lite. It also surpasses an oracle best-of-five baseline at substantially lower cost, showing that the gains come from better feedback rather than additional test-time sampling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑