快速失败,智能重启:面向软件工程智能体任务的早期失败预测与重启
Fail-Fast, Restart-Smart: Early Failure Prediction and Restart for SWE Agentic Tasks
另 1 家 · 查看机构详情
- Singapore Management University(新加坡管理大学)
- University of Alberta(阿尔伯塔大学)
- Nanjing University of Science and Technology(南京理工大学)
- Tel Aviv University(特拉维夫大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出FailFast-RestartSmart两阶段控制器,可在5%假阳性率下节省14.6%-20.4%执行token,在25%假阳性率下将Qwen3.6-27B解决率从66.6%提至71.8%。
中文摘要 AI 辅助
软件工程(SWE)智能体通过长轨迹解决仓库级问题,随着上下文积累,计算成本不断上升。失败的运行往往持续时间更长,表现出冗余探索或循环,这表明部分失败可在完成前被检测到。然而,提前终止可能会中断原本会成功的轨迹;相反,未成功的轨迹可能仍包含有用的仓库编辑。本文提出FailFast-RestartSmart,一种针对单个活跃轨迹的两阶段控制器。FailFast是一个轻量级0.6B规模的监控器,使用终端和密集的“失败转通过”监督进行训练,可从可观测前缀预测失败,无需策略逻辑或隐藏状态。发出警报后,RestartSmart会启动新的相同策略rollout,且不包含先前的提示历史,并将中断的仓库差异作为可选叠加项,供智能体检查、应用或丢弃。在SWE-bench Verified上,仅基于Qwen3.6-27B轨迹训练的监控器可迁移至其他三个策略(包括闭API模型),在5%的目标假阳性率下节省14.6%-20.4%的执行token;针对Qwen3.6-27B,其20.4%的节省幅度超过每步AgentStop适配实现的12.5%。在25%的目标假阳性率下,RestartSmart将Qwen3.6-27B的解决率从66.6%提升至71.8%,而冷重启仅达到66.8%。这些结果共同支持了结合早期终止与相同策略顺序恢复的方案。
英文摘要
Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context accumulates. Failed runs tend to be longer and exhibit redundant exploration or looping, suggesting that some failures may be detectable before completion. Early termination, however, risks interrupting trajectories that would otherwise succeed; conversely, an unsuccessful trajectory may still contain useful repository edits. We present FailFast-RestartSmart, a two-stage controller for a single active trajectory. FailFast is a lightweight 0.6B monitor trained with terminal and dense fail-to-pass supervision to predict failure from observable prefixes without policy logits or hidden states. Upon an alarm, RestartSmart launches a fresh same-policy rollout without prior prompt history and offers the interrupted repository diff as an optional overlay that the agent may inspect, apply, or discard. On SWE-bench Verified, a monitor trained solely on Qwen3.6-27B trajectories transfers to three other policies, including a closed-API model, and saves 14.6%-20.4% of execution tokens at a target 5% false-positive rate; on Qwen3.6-27B, its 20.4% saving exceeds the 12.5% achieved by our per-step AgentStop adaptation. At a target 25% false-positive rate, RestartSmart raises Qwen3.6-27B resolution from 66.6% to 71.8%, whereas cold restart reaches only 66.8%. Together, these results support early termination with sequential same-policy recovery.