arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

相同的胜者,不同的成功率:评估LLM智能体如何从失败中恢复

Same Winners, Different Success Rates: Evaluating How LLM Agents Recover from Failures

Dong Xu, Zhangfan Yang, Jiantao Wu, Shipeng Zhang, Zexuan Zhu, Jiangqiang Li, Jun Zhang, Junkai Ji

arXiv 2609.34215首次发表:更新:

发表机构

Shenzhen University; EasternDawn; University of Nottingham Ningbo(深圳大学; 东方黎明; 宁波诺丁汉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示LLM智能体恢复评估中集合一致性指标的序数模糊性缺陷,提出合并成功率等四个诊断量,并通过理论证明和实验验证其有效性。

AI 中文摘要

评估LLM智能体如何从任务中途失败中恢复,是部署可靠智能体系统的核心问题。现有的基于检查点的基准通过比较独立运行中选择哪个动作作为最佳动作来衡量恢复能力,这一指标被称为集合一致性。然而,集合一致性是一种纯粹的序数度量,它记录哪个动作获胜,而不反映绝对性能水平。当所有动作都失败时,它们在零奖励处并列,独立运行以高概率产生相同的并列集合,从而造成一种掩盖接近零恢复成功率的稳定性假象。我们通过集合路径对称性结果形式化了这一局限性,证明对于等成本伯努利动作,成功概率(0.9, 0.8)和(0.2, 0.1)在任意样本量下都会产生相同的最佳动作集合分布。仅基于哪个动作获胜的任何过程都无法区分这两种情形。我们进一步证明,在有限时间内证明确切的总体并列是不可能的,并且结果到检查点的分配所携带的信息超出了边际结果分布。合并成功概率是解决序数模糊性的缺失标量。在864个冻结的RecoveryBench情节和总计3,456个响应的两个规划队列上的实验证实了理论预测。一致性和保留质量可能朝相反方向移动,并且置换检查点到动作的绑定会改变8%至13%的单元级结论。基于这些发现,我们建议报告四个诊断量(一致性、全零比例、保留成功率和合并成功率),这些量无需额外数据收集即可暴露这种失败模式。

英文摘要

Evaluating how LLM agents recover from mid-task failures is central to deploying reliable agentic systems. Existing checkpoint-based benchmarks measure recovery by comparing which action is selected as best across independent runs, a quantity known as set agreement. However, set agreement is a purely ordinal measure that records which action wins without reflecting the absolute level of performance. When all actions fail, they tie at zero reward, and independent runs produce the same tied set with high probability, creating an illusion of stability that masks near-zero recovery success. We formalize this limitation through a set-path symmetry result, proving that for equal-cost Bernoulli actions the success probabilities (0.9, 0.8) and (0.2, 0.1) yield identical best-action-set distributions at every sample size. No procedure based solely on which action wins can distinguish these two regimes. We further prove that certifying exact population ties is impossible in finite time, and that the assignment of outcomes to checkpoints carries information beyond marginal outcome distributions. The pooled success probability is the missing scalar that resolves the ordinal ambiguity. Experiments on 864 frozen RecoveryBench episodes and two planning cohorts totaling 3,456 responses confirm the theoretical predictions. Agreement and held-out quality can move in opposite directions, and permuting checkpoint-to-action bindings changes 8 to 13 percent of cell-level conclusions. Based on these findings, we propose reporting four diagnostic quantities (agreement, all-zero fraction, held-out success, and pooled success) that expose this failure mode with no additional data collection.

Comments44 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑