AI 中文总结
针对工具使用型AI智能体,提出UndoBench基准,通过反事实配对试验解耦任务能力与故障恢复,发现名义完成率高但条件恢复成功率显著下降,揭示阶段依赖的恢复漏洞。
AI 中文摘要
工具使用型AI智能体正越来越多地部署在企业软件系统中,然而广泛使用的基准主要评估名义任务完成情况,将基线规划能力与操作故障恢复混为一谈。我们引入了UndoBench,这是一个涵盖8个企业领域、36个基础工作流和36个故障场景的基准,通过相同种子下的反事实配对试验,结合线路级效果历史和环境状态预言机,将任务能力与恢复能力解耦。在冻结的丢失确认研究中,针对12个保留的TEST工作流,使用两个开放权重模型、两个框架和三种恢复范式(共5,760次执行/2,880次配对试验),名义能力达到83.54%,而条件恢复成功率(CRSR)降至46.72%,其中朴素重试在53.33%的试验中产生了重复的外部效果。扩展到商业API模型也重现了这种能力与恢复的分离。跨互补执行边界的评估表明,恢复具有阶段依赖性:在突变之前,各方法在能力试验中表现相似且无重复效果;在部分突变期间,朴素重试、每次调用幂等性和零权限日志记录在评估的复合工作流上均失效;在提交后但确认前,验证和服务器端幂等性显著提高了安全性。这些发现表明,仅评估名义完成情况会掩盖自主智能体中关键的、阶段依赖的恢复漏洞。
英文摘要
Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating baseline planning competence with operational fault recovery. We introduce UndoBench, a benchmark spanning 36 base workflows and 36 fault scenarios across 8 enterprise domains, decoupling task competence from recovery capability via counterfactual paired trials under identical seeds alongside wire-level effect-history and environment-state oracles. On 12 held-out TEST workflows across two open-weight models, two frameworks, and three recovery paradigms (5,760 executions / 2,880 paired trials) in the frozen lost-acknowledgment study, nominal competence reached 83.54% while conditional recovery success rate (CRSR) fell to 46.72%, with naive retry producing duplicate external effects in 53.33% of trials. Extensions to commercial API models reproduced this competence-recovery separation. Evaluations across complementary execution boundaries show that recovery is phase-dependent: before mutation, methods perform similarly without duplicate effects among capable trials; during partial mutation, naive retry, per-call idempotency, and zero-privilege journaling collapse on the evaluated composite workflows; after commit but before acknowledgment, verification and server-side idempotency substantially improve safety. These findings demonstrate that evaluating nominal completion alone masks critical, phase-dependent recovery vulnerabilities in autonomous agents.
Comments18 pages, 6 figures. Code and benchmark available at https://github.com/tradertanmay/undobench