在软件问题解决的大语言模型智能体中耦合规划与情景记忆
Coupling Planning with Episodic Memory in LLM Agents for Software Issue Resolution
浏览论文内容
中文总结 AI 辅助
本文提出PMCoder智能体,通过双向耦合分层阶段规划器与情景记忆解决软件问题,在SWE-bench Verified等数据集上显著提升问题解决能力,且性能可迁移至其他场景。
中文摘要 AI 辅助
用大语言模型(LLM)智能体解决真实软件问题是一个漫长的修复流程,通常包含数十到数百个步骤,涵盖探索、假设、实现和验证环节。解决效果既取决于基础模型的局部推理能力,也取决于智能体维持动态规划并跨阶段记忆观测结果的能力。现有的仓库级智能体通常单独强化规划或记忆能力,导致长轨迹易受陈旧证据、重复失败编辑的影响,且验证依赖智能体自身声明而非执行证据。本文提出PMCoder,一种将分层阶段规划器与情景记忆耦合的问题解决智能体,二者为双向耦合:当前计划阶段为记忆检索提供条件,而记忆衍生的轨迹统计信息用于停滞检测与重新规划。若可用,问题复现裁决会将验证进度建立在执行证据而非自报告完成情况上。在SWE-bench Verified数据集上,PMCoder平均多解决25个案例(提升5.0个百分点),优于匹配工具的基线,即使在未触发复现闸门的场景中,该优势仍持续存在。进一步在Verified-500上的评估显示,Claude Haiku 4.5、DeepSeek-V4-Flash及OpenHands移植版本均呈现相同正向趋势,至少多解决14个案例(提升2.8个百分点)。此外,对TerminalWorld官方样本的评估表明,该规划-记忆基底可迁移至软件问题报告之外的场景。消融实验与轨迹分析揭示了优势来源:规划与记忆的耦合性能优于任一组件单独作用,并减少重复失败动作、空补丁退出及上下文窗口耗尽的情况。
英文摘要
Resolving a real software issue with a large language model (LLM) agent is a long repair episode, often tens to hundreds of steps spanning exploration, hypothesis, implementation, and verification. Success depends on both the base model's local reasoning and the agent's ability to maintain an evolving plan and remember observations across phases. Existing repository-level agents typically strengthen planning or memory in isolation, leaving long trajectories vulnerable to stale evidence, repeated failed edits, and verification inferred from the agent's own claims instead of execution evidence. We present PMCoder, an issue-resolution agent that couples a hierarchical phase planner with episodic memory. The coupling is bidirectional: the current plan phase conditions memory retrieval, while memory-derived trajectory statistics inform stuck detection and replanning. When available, issue-reproduction verdicts ground verification progress in execution evidence rather than self-reported completion. On SWE-bench Verified, PMCoder resolves an average of $25$ more cases ($+5.0$pp) than a harness-matched baseline, with gains persisting even where the reproduction gate never fires. Further Verified-500 evaluations show the same positive direction across Claude Haiku 4.5, DeepSeek-V4-Flash, and an OpenHands port, with at least $14$ additional resolved cases ($+2.8$pp). Separately, evaluation on TerminalWorld's official sample suggests that the plan-memory substrate transfers beyond issue reports. Ablation and trajectory analyses show where the gains come from: coupling planning and memory outperforms either component alone and reduces repeated failed actions, empty-patch exits, and context-window exhaustion.
发表机构
- Vanderbilt University(范德堡大学)
机构由 AI 辅助整理,请以论文原文为准。