arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Scrouting:通过先探查仓库实现编码智能体的成本感知路由

Scrouting: Cost-Aware Routing of Coding Agents by Scouting the Repository First

Ishaan Bhola, Adithyan Krishnan, Mukunda NS

arXiv 2608.04804首次发表:更新:

发表机构

SuperAGI Research(SuperAGI研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SuperScout,先探查仓库生成经沙箱验证的结构化交接内容,再路由任务至四个修复器,在SWE-bench Pro上以约五分之一成本达到最优单模型的解决率。

AI 中文摘要

前沿语言模型可解决仓库级软件问题,但每次尝试成本高昂,现有路由器仅根据问题文本选择模型。本文提出SuperScout,其在探查仓库后再进行路由:一个7B规模的探查器SuperScout-7B先探索仓库,生成结构化交接内容,该内容的复现声明经沙箱验证,交付前会剔除虚假声明。探查器的隐藏状态与任务文本随后输入基于简历的路由器,将任务分派给四个前沿修复器之一,新增修复器无需重新训练。在SWE-bench Pro的完整Python子集(266个任务)及基准官方上限预算层级下,SuperScout的解决率与最优单模型持平(SuperScout为266个任务中解决159个,最优模型为158个),但每个解决任务的总成本仅为最优单模型的约五分之一,且报告配置优于随机流量拆分基线。无路由器的 ablation(消融)实验中,始终使用带交接内容的最便宜修复器,在该基准上与带路由的系统表现相当,故结果由交接内容而非路由决策承载。配对校准研究指出机制:交接内容似是重新分配而非增加解决能力,提升了三个较便宜修复器的表现,同时略微削弱最强修复器的表现(尽管N=99时各修复器的影响仅具方向性);探查器的隐藏状态在校准标签上改善了成本路由,而交接内容自身文本则无此效果。探查器的计算开销每个任务不到半美分的GPU时间。

英文摘要

Frontier language models can resolve repository-level software issues, but each attempt is expensive, and existing routers select a model from the issue text alone. We present SuperScout, which routes after scouting the repository: a 7B searcher, SuperScout-7B, first explores the repository and produces a structured handoff whose reproduction claims are sandbox-verified, with false claims stripped before delivery. The searcher's hidden states, together with the task text, then feed a resume-based router that dispatches the task to one of four frontier fixers. Adding a new fixer requires no retraining. On the full Python slice of SWE-bench Pro (266 tasks) under the benchmark's official capped budget tier, SuperScout matches the best single model's solve rate (159 of 266 for SuperScout, 158 for the best model) at about a fifth of the total cost per solve, and the reported configuration sits above the random traffic-splitting baseline. A no-router ablation, always the cheapest fixer with the handoff, ties the routed system on this benchmark, so the handoff rather than the routing decision carries the result. A paired calibration study points to the mechanism: the handoff appears to redistribute rather than add solving ability, lifting the three cheaper fixers while slightly hurting the strongest, though at $N=99$ the per-fixer effects are directional only; the searcher's hidden states improve cost routing on the calibration labels while the handoff's own text does not. The searcher's compute adds less than half a cent of GPU time per task.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑