CodeRescue:用于编码智能体的预算校准恢复路由
CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents
浏览论文内容
中文总结 AI 辅助
研究编码智能体失败后的预算部署问题,核心方法是将其制定为异构动作上的恢复路由并训练监督路由器,添加CRC层实现预算变化下的成本控制,主要贡献是校准前沿在多个方面优于基线,节省恢复成本。
中文摘要 AI 辅助
编码智能体越来越多地在可执行环境中运行,失败尝试会产生可操作的反馈而非仅仅是错误答案。现有成本感知系统通常将此类失败视为级联决策。然而在编码中,执行反馈能使廉价模型恢复有价值,引出预算部署问题。我们将失败后决策制定为异构动作上的恢复路由,并从执行展开中训练监督路由器。为使路由器在变化预算下可用,添加共形风险控制(CRC)层,其无需重新训练就能选择部署时成本惩罚并提供边际预期成本控制。在五个编码基准的保留失败案例中,廉价恢复和升级呈现互补成功模式。校准前沿优于固定动作、仅提示路由器和二元级联基线;在主要的GPT - 5.4 - nano/GPT - 5.4设置中,一个CRC校准前沿点超过始终升级的解决率,同时使用其平均恢复成本的35%。代码可在指定网址获取。
英文摘要
Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer. Existing cost-aware systems typically treat such failures as cascade decisions: try a cheap model first, then escalate hard cases to a stronger and more expensive model. In coding, however, execution feedback can also make further cheap-model recovery worthwhile, raising a budgeted deployment question: when should an agent spend more cheap compute, and when should it escalate? We formulate this post-failure decision as recovery routing over heterogeneous actions and train a supervised router from execution rollouts. To make the same router usable under changing budgets, we add a Conformal Risk Control (CRC) layer that selects a deployment-time cost penalty without retraining and provides marginal expected-cost control under exchangeability. Across held-out failures from five coding benchmarks, cheap recovery and escalation exhibit complementary success patterns. The calibrated frontier improves over fixed actions, prompt-only routers, and a binary cascade baseline; in the main GPT-5.4-nano/GPT-5.4 setting, one CRC-calibrated frontier point exceeds always-escalate solve rate while using 35% of its mean recovery cost. Code is available at https://github.com/Qijia-He/agent-budget-control.
发表机构
- University of Washington(华盛顿大学)
- New York University(纽约大学)
- ByteDance(字节跳动)
- Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。