arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过分层轨迹抽象复用编码智能体的过往修复工作

Reusing Past Repairs Through Hierarchical Trajectory Abstraction for Coding Agents

Yisen Xu, Jiayuan Zhou, Ruiqi Pan, Tse-Hsun Chen

arXiv 2607.29658首次发表:更新:

AI 中文总结

本研究提出STAIR框架,通过分层轨迹抽象复用编码智能体过往修复经验,在SWE-bench Verified上显著提升修复智能体的Pass@1指标,且计划可跨智能体泛化。

AI 中文摘要

尽管大语言模型驱动的修复智能体能够处理复杂的仓库级问题,但它们会将每个问题视为独立个体,丢弃从过往修复中积累的过程性知识。本文提出STAIR框架,该框架可将历史修复轨迹转换为分层的可复用计划,用于指导未来的修复工作。每条过往轨迹被转化为从细粒度诊断操作到高层修复策略的多级树结构,编码不同粒度的经验。当新问题出现时,STAIR会从多个抽象层级中选择相关的计划节点,将其定制为适用于特定问题的可执行计划,并通过提示提供给智能体。在SWE-bench Verified基准测试中,结合STAIR的Lingxi框架使用MiniMax M2.5模型达到81.2%的Pass@1指标,使用GPT-5模型达到79.2%。生成的计划还可在不同智能体间泛化:无需任何代码修改,它们将结构不同的智能体mini-SWE-agent v2的Pass@1从75.8%提升至81.0%。 ablation实验进一步表明,混合多个抽象层级的效果优于任何单一层级,而未抽象的原始轨迹的迁移效果则差得多。

英文摘要

Although LLM-driven repair agents can tackle complex, repository-level issues, they treat every issue independently and discard the procedural knowledge accumulated from previous repairs. We introduce STAIR, a framework that converts historical repair trajectories into hierarchical, reusable plans that can be adapted to steer future repairs. Each past trajectory is transformed into a multi-level tree that ranges from fine-grained diagnostic actions to high-level repair strategies, encoding experience at several granularities. When a new issue arrives, STAIR selects relevant plan nodes from multiple abstraction levels, tailors them into executable, issue-specific plans, and supplies them to the agent through its prompt. On SWE-bench Verified, STAIR integrated with Lingxi reaches 81.2% Pass@1 using MiniMax M2.5 and 79.2% using GPT-5. The generated plans also generalize across agents: without any code change, they lift the Pass@1 of a structurally different agent, mini-SWE-agent v2, from 75.8% to 81.0%. Ablation experiments further show that mixing multiple abstraction levels surpasses any single level and that raw, unabstracted trajectories transfer substantially worse.

Comments10 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑