在智能体失败预测中区分任务难度与运行级失败
Disentangling Task Difficulty from Run-Level Failure in Agent Failure Prediction
浏览论文内容
中文总结 AI 辅助
本研究揭示LLM智能体失败预测的高性能主要源于任务难度而非运行级失败,提出区分两者,并证明任务级分配优于中止策略,早期停止需高任务内AUROC才有效。
中文摘要 AI 辅助
预测LLM智能体是否会失败已成为在执行过程中支持干预的一个有前景的方向。近期方法报告了强大的预测性能,AUROC值通常在0.85至0.94之间。然而,预测器通常通过汇集来自许多任务的数据进行训练。我们假设这种性能的一部分来自于识别某些任务比其他任务更难,而不是检测特定运行是否正走向失败。这一区别很重要,因为任务级难度支持关于在何处分配计算资源的决策,而运行级预测则需要决定是否干预正在进行的轨迹。我们研究了同一智能体对同一任务进行重复尝试的基准,并将跨单元比较与同一模型-任务单元的成功和失败运行之间的比较分开。在评估的语料库中,超过99.93%的构成汇集AUROC的正负对是跨单元的。因此,从未观察当前运行的预测器可以实现高汇集性能,包括AUROC高达0.945的难度神谕,而在任务内保持随机水平。运行级早期判别在轨迹预测器、发布的监控器和隐藏状态探针中持续较弱,尽管它在执行后期有所改善,并且对于较弱的智能体更强。在固定的令牌预算下,任务级分配优于仅中止策略,而早期停止仅在任务内AUROC达到约0.84-0.93时才有益,远高于早期监控器观察到的0.50-0.55范围。这些结果表明,失败预测不仅应通过结果准确性来评估,还应通过捕获的信号是否支持预期的部署决策来评估。
英文摘要
Predicting whether an LLM agent will fail has emerged as a promising direction for supporting intervention during execution. Recent approaches report strong predictive performance, often with AUROC values between 0.85 and 0.94. However, predictors are typically trained by pooling runs from many tasks. We hypothesize that part of this performance comes from recognizing that some tasks are harder than others, rather than detecting whether a particular run is heading toward failure. This distinction matters because task-level difficulty supports decisions about where to allocate computation, while run-level prediction is needed to decide whether to intervene in an ongoing trajectory. We study benchmarks with repeated attempts of the same task by the same agent and separate cross-unit comparisons from comparisons between successful and failed runs of the same model-task unit. Across the evaluated corpora, more than 99.93% of the positive-negative pairs underlying pooled AUROC are cross-unit. Accordingly, predictors that never observe the current run can achieve high pooled performance, including a difficulty oracle with AUROC up to 0.945, while remaining at chance within task. Early run-level discrimination is consistently weak across trajectory predictors, released monitors, and hidden-state probes, although it improves later in execution and is stronger for weaker agents. Under fixed token budgets, task-level allocation outperforms abort-only strategies, while early stopping becomes beneficial only when within-task AUROC reaches about 0.84-0.93, far above the 0.50-0.55 range observed for early monitors. These results show that failure prediction should be evaluated not only by outcome accuracy, but by whether the captured signal supports the intended deployment decision.
发表机构
- North Carolina State University(北卡罗来纳州立大学)
机构由 AI 辅助整理,请以论文原文为准。