arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2604.11716cs.AIcs.CL

SWE-AGILE:一种用于高效管理动态推理上下文的软件代理框架

SWE-AGILE: A Software Agent Framework for Efficiently Managing Dynamic Reasoning Context

  • Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

Shuquan Lian, Juncheng Liu, Yazhe Chen, Yuhong Chen, Hui Li

更新

AI总结:

SWE-AGILE通过动态推理上下文策略平衡推理深度与效率,实现高效多轮软件工程任务处理,实验证明其在SWE-Bench-Verified上优于7B-8B模型。

AI中文摘要:

在自主软件工程(SWE)中,传统ReAct风格方法缺乏深度分析和处理复杂边缘情况所需的显式系统2推理。尽管最近的推理模型展示了扩展链式思维(CoT)的潜力,但将其应用于多轮SWE任务时会面临根本性困境:保留完整推理历史会导致上下文爆炸和『迷失中间』退化,而丢弃则迫使代理在每一步都重复推理。为解决这些挑战,我们提出了SWE-AGILE,一种新的软件代理框架,旨在弥合推理深度、效率和上下文限制之间的差距。SWE-AGILE引入了动态推理上下文策略,维护一个『滑动窗口』的详细推理以保持即时连续性,防止重复分析,同时将历史推理内容压缩成简洁的推理摘要。实验证明,SWE-AGILE在仅使用2.2k轨迹和896个任务的情况下,为7B-8B模型设定了新的标准。代码可在https://github.com/KDEGroup/SWE-AGILE获取。

英文摘要:

Prior representative ReAct-style approaches in autonomous Software Engineering (SWE) typically lack the explicit System-2 reasoning required for deep analysis and handling complex edge cases. While recent reasoning models demonstrate the potential of extended Chain-of-Thought (CoT), applying them to the multi-turn SWE task creates a fundamental dilemma: retaining full reasoning history leads to context explosion and ``Lost-in-the-Middle'' degradation, while discarding it would force the agent to redundantly re-reason at every step. To address these challenges, we propose SWE-AGILE, a novel software agent framework designed to bridge the gap between reasoning depth, efficiency, and context constraints. SWE-AGILE introduces a Dynamic Reasoning Context strategy, maintaining a ``sliding window'' of detailed reasoning for immediate continuity to prevent redundant re-analyzing, while compressing historical reasoning content into concise Reasoning Digests. Empirically, SWE-AGILE sets a new standard for 7B-8B models on SWE-Bench-Verified using only 2.2k trajectories and 896 tasks. Code is available at https://github.com/KDEGroup/SWE-AGILE.

↑