AI 中文总结
SWE-MILE通过异步势诱导里程碑信用分配,利用工作流运行时信号为长程软件工程智能体提供细粒度过程监督,显著提升其性能。
AI 中文摘要
使用可验证奖励的强化学习(RLVR)训练的长程软件工程(SWE)智能体通常仅接收终端结果监督,这使得难以区分有效行动与冗余探索或功能回退。我们提出SWE-MILE,一种异步势诱导里程碑信用分配框架,该框架从工作流运行时中提取细粒度过程监督,无需辅助奖励模型或外部评估器。SWE-MILE分别将任务相关文件暴露度和测试状态对齐度量化为导航势和验证势。这些势的差异将里程碑进展和回退归因于各个行动,而折扣后向信用则将监督传播至前置步骤。为高效获取中间验证状态,SWE-MILE进一步引入异步影子探测,在隔离沙箱中重放修改仓库的行动,并与智能体的主交互并行运行验证,从而在很大程度上隐藏验证延迟。所得过程信用增强终端结果优势并提供信息丰富的学习信号。在两个代表性长程SWE任务上的实验表明,智能体性能有显著提升,凸显了工作流运行时信号作为长程SWE智能体过程监督的实际来源。
英文摘要
Long-horizon software engineering (SWE) agents trained with reinforcement learning with verifiable rewards (RLVR) typically receive only terminal outcome supervision, making it difficult to distinguish productive actions from redundant exploration or functional regressions. We propose SWE-MILE, an asynchronous potential-induced milestone credit assignment framework that derives fine-grained process supervision from workflow runtime, without auxiliary reward models or external evaluators. SWE-MILE quantifies task-relevant file exposure and test-state alignment as navigation and verification potentials, respectively. Differences in these potentials attribute milestone progress and regressions to individual actions, while discounted backward credit propagates supervision to preceding steps. To efficiently acquire intermediate verification states, SWE-MILE further introduces asynchronous shadow probing, which replays repository-changing actions in an isolated sandbox and runs verification in parallel with the agent's primary interaction, largely hiding verification latency. The resulting process credit augments terminal outcome advantages and provides informative learning signals. Experiments on two representative long-horizon SWE tasks demonstrate substantial improvements in agent performance, highlighting workflow runtime signals as a practical source of process supervision for long-horizon SWE agents.
Comments23 pages, 6 figures