AI 中文总结
本文提出三轴分类法、TC-ECE评估协议,通过实验揭示LLM智能体不确定性源于多轮交互,单一置信度分数不足,步骤级校准不保证轨迹级校准。
AI 中文摘要
大型语言模型(LLM)不再仅用于单轮对话,而是越来越多地作为智能体(agent)来规划、调用工具、检索证据、维护记忆,并在长时间跨度内进行交互,通常还通过多轮对话与其他智能体协作。因此,知道何时信任智能体系统是安全部署的前提。然而,现有的LLM不确定性量化工作几乎完全针对单轮问答构建。本文认为,错误和不确定性源于多轮对话、环境和工具,而非单轮问答场景。这种不确定性出现较晚,并被合并为一个过于粗糙的单一分数,无法代表不可靠性。我们用一个三轴分类法组织文献:(1)不确定性是什么,(2)如何估计不确定性,(3)不确定性在智能体流程中何处产生。我们研究了步骤级和轨迹级校准,并用一个简单的反例表明前者并不意味着后者。在四个模型和最多50步预算的真实智能体轨迹上的实验表明,所提出的指标和报告协议(轨迹检查点期望校准误差,TC-ECE)可以计算,并且步骤错误沿轨迹耦合。我们发现,来自智能体自身响应的置信度估计并不始终优于简单基线。实验还表明,将所有轨迹平均在一起会掩盖后期阶段的过度自信,而按不同时间跨度分析结果时,这种过度自信会显现。简言之,本文指出了智能体系统流程中不确定性的来源,如何教会智能体知道自己何时出错,以及为什么一个置信度数字是不够的。
英文摘要
Large language models (LLMs) are no longer deployed only for single-turn conversation but increasingly act as agents that plan, call tools, retrieve evidence, maintain memory, and interact over long horizons, often together with other agents through multi-turn conversations. Therefore, knowing when to trust the agentic system is a prerequisite for safe deployment. However, existing work on quantifying uncertainty for LLMs was built almost entirely for single-turn question answering. This paper argues that errors and uncertainty arise from multi-turn conversations, environments, and tools rather than from a single-turn question answering setting. It comes late, however, and is compounded in a single score that is too coarse to represent the unreliability. We organize the literature with a three-axis taxonomy, (1) what the uncertainty is, (2) how it is estimated, and (3) where uncertainty arises during an agent pipeline. We investigate step-level and trajectory-level calibration and show with a simple counterexample that the first does not imply the second. Experiments on real agent traces across four models and up to a 50-step budget show that the proposed metric and reporting protocol (Trajectory-Checkpoint Expected Calibration Error, TC-ECE) can be computed and that step errors are coupled along a trajectory. We find that confidence estimates from the agent's own responses do not consistently outperform a simple baseline. The experiments also show that averaging all trajectories together can hide overconfidence at later stages, which becomes visible when results are analyzed across different horizons. In simpler terms, this paper identifies where the uncertainty comes from in the agentic system pipeline, how to teach agents to know when they are wrong, and why one confidence number is not enough.