发表机构
Florida International University(佛罗里达国际大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出反事实检查点优势度量方法,通过对比检查点与跳过分支的恢复成本,发现首个检查点收益显著而后续检查点可能为负,并指出恢复是重推导而非重放,为检查点放置策略设定标准。
AI 中文摘要
智能体检查点系统决定哪些状态与恢复相关、如何对其进行快照,以及回滚是否可接受。但没有任何系统决定其所暴露的安全边界中哪些值得物化。我们将此表述为反事实检查点优势,即通过对候选状态设置检查点而非跳过它而获得的未来恢复成本的降低,并通过在匹配的模型、工具、验证器和停止条件下,将检查点分支(CP分支)和跳过分支(SKIP分支)驱动到相同的逻辑故障并恢复两者来度量该优势。在包含12个SWE-bench Verified任务和106个真实恢复分支的冻结试点中,设置检查点每任务节省49.4秒,且该数字分解为相差两个数量级的两个区间。第一个检查点在156.5秒受保护工作上返回100.0秒,转换率为0.64;一步之后的第二个检查点在41.8秒上返回-1.1秒,转换率为-0.03。恢复是重新推导而非重放,因此保留的工作量是节省工作量的不良指标,且经典的已用工作量规则以完全名义成本错误定价第二个检查点。我们确定了放置可能带来收益的位置,并为放置策略设定了必须达到的标准。
英文摘要
Agent checkpoint systems decide what state is recovery-relevant, how to snapshot it, and whether rollback is admissible. None decides which of the safe boundaries they expose are worth materializing. We formulate this as counterfactual checkpoint advantage, the reduction in future recovery cost obtained by checkpointing a candidate rather than skipping it, and measure it by driving a CP branch and a SKIP branch to the same logical failure and recovering both under matched model, tool, verifier, and stopping conditions. On a frozen pilot of 12 SWE-bench Verified tasks and 106 real recovery branches, checkpointing saves 49.4 s per task, and that figure resolves into two regimes two orders of magnitude apart. The first checkpoint returns 100.0 s on 156.5 s of protected work, a conversion of 0.64; a second one step later returns-1.1 s on 41.8 s, a conversion of -0.03. Recovery is a re-derivation rather than a replay, so preserved work is a poor guide to saved work, and the classical elapsed-work rule misprices the second checkpoint by its full nominal cost. We identify where placement can pay, and set the bar a placement policy must clear.
Comments16 pages