发表机构
University of Illinois Urbana-Champaign; IBM(伊利诺伊大学厄巴纳-香槟分校; 国际商业机器公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出LivePlan,通过解耦判断与建议的设计,在SWE-agent基础上实现编程智能体的在线监控与纠正,在SWE-bench数据集上显著提升问题解决率,且成本增量极小。
AI 中文摘要
修复大规模项目中的GitHub问题是一项长期任务,尤其当修复需要跨多个位置修改,或问题描述缺乏定位与修复所需信息时。因此,智能体的行动轨迹漫长,易出现低效和错误:它们会偏离预期计划、重复失败的动作,或在未生成有效补丁的情况下终止。本文提出LivePlan,用于实时监控、检测和纠正此类行为低效与偏差。LivePlan将判断与建议解耦:确定性的基于规则的监控器检查轨迹上的通用信号以检测问题,无需调用大语言模型(LLM),仅在检测到问题时才咨询顾问LLM以获取高层级的下一步纠正方案。该设计避免了先前方法中误导性的重新规划和高成本干预。我们在SWE-agent基础上实现LivePlan,使用5个LLM(3个作为执行智能体,2个作为顾问)在SWE-bench Verified和SWE-bench Pro上进行评估。与原始SWE-agent相比,LivePlan显著提升了问题解决率,实现了最高15.2%的稳定提升(平均9.9%),而每个实例仅产生额外0.08美元的成本。额外的解决方案集中在中等难度和困难实例上。LivePlan在解决率上始终优于替代方法,对已成功运行的回归影响极小,且在基准方法无法解决的问题上取得了新的成功。
英文摘要
Fixing GitHub issues in large-scale projects is a long-horizon task, especially when a fix requires changes across multiple locations or the issue description lacks the information needed to localize and repair it. As a result, agents traverse long trajectories that are prone to inefficiency and error: they drift away from their intended plan, repeat failed actions, or terminate without a working patch. This paper proposes LivePlan to monitor, detect, and correct such behavioral inefficiencies and drifts in real time. LivePlan decouples judging from advising: a deterministic, rule-based monitor examines general signals over the trajectory to detect issues without invoking an LLM, and only when an issue is detected does it consult an advisor LLM for a high-level, next-step correction. This design avoids the misleading re-planning and costly interventions of prior approaches. We implement LivePlan on top of SWE-agent and evaluate it using five LLMs (three as executor agents and two as advisors) across SWE-bench Verified and SWE-bench Pro. Compared to vanilla SWE-agent, LivePlan notably improves issue resolution rates, achieving consistent gains of up to 15.2% (average: 9.9%), while incurring only an additional cost of $0.08 per instance. The additional solutions concentrate on medium and hard instances. LivePlan consistently outperforms alternative approaches in resolution rate, with minimal regression on already successful runs and new successes on problems that no baseline solves.