发表机构
Zhejiang University; Ant Group(浙江大学; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现长视界智能体的早期不确定性无法预测失败,仅最后一步的语言置信度可可靠区分失败,建议用最后一步置信度决定是否重启,其效果优于轨迹内干预。
AI 中文摘要
早期失败预测对长视界智能体至关重要,因为它能实现及时干预,降低推理和工具使用成本。不确定性量化(如语言置信度和困惑度)为检测智能体失败提供了有前景的方法,但尚未探究这些信号在长视界执行的中间阶段是否仍保留其判别能力。我们在深度研究任务上评估主流不确定性信号,发现语言置信度在轨迹完成时可可靠区分失败,平均AUROC达0.85;而所有评估信号在执行早期的预测价值有限,在轨迹进度达50%时,无一超过平均AUROC 0.60。我们揭示了解释这一差距的潜在机制:路径切换,即智能体在轨迹中频繁放弃当前搜索方向,打破了早期信号与最终结果的关联。这些发现挑战了中间不确定性可可靠指导早期干预的假设,也为深度研究场景中的智能体管控提供了实用建议:使用最后一步的置信度决定是否重启,我们的实验表明该方法比轨迹内干预更有效。
英文摘要
Early failure prediction is important for long-horizon agents, as it enables timely intervention and can reduce inference and tool-use costs. Uncertainty quantification, such as verbal confidence and perplexity, offers a promising approach to detecting agent failures; however, it has not been explored whether these signals retain their discriminative power during the intermediate stages of long-horizon execution. We evaluate mainstream uncertainty signals on deep-research tasks and find that verbal confidence reliably distinguishes failures at trajectory completion, achieving a mean AUROC of 0.85, whereas all evaluated signals offer limited predictive value earlier in execution, with none exceeding a mean AUROC of 0.60 at 50% trajectory progress. We identify an underlying mechanism explaining this gap: path switching, where agents frequently abandon their current search direction in-trajectory, breaking the link between early signal and final outcome. These findings challenge the assumption that intermediate uncertainty can reliably guide early intervention. They also motivate a practical recommendation for agent harnesses in deep-research settings: use final-step confidence to decide whether to restart, an approach that our experiments find more effective than in-trajectory intervention.
CommentsAccepted to the Main Conference of the 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)