发表机构
University of California, Santa Cruz; Megagon Labs(加州大学圣克鲁兹分校; Megagon 实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出置信推理图(CRG),一种无需特权模型访问或训练数据的推理时框架,通过将任务声明分解为子声明并聚合置信度,在多个基准上实现更优的校准置信度与风险感知决策。
AI 中文摘要
在具有重要后果的领域中使用LLM智能体时,做出关于是否信任其输出或进行干预的明智决策,需要对智能体的成功概率进行校准的置信估计。智能体的置信估计之所以困难,是因为关于成功的证据分散在智能体轨迹中异构且相互依赖的各个步骤中。实际的智能体部署带来了进一步的挑战:前沿LLM通常对内部信号的访问有限,智能体 rollout 成本高昂,且训练数据可能不可用或很快过时。为应对这些挑战,我们引入了置信推理图(CRGs),这是一种推理时框架,可从单条轨迹估计智能体完成其任务的概率,无需特权模型访问或训练数据。CRG并非将执行压缩为单一整体判断,而是从智能体已完成其任务的声明出发,将其分解为基于轨迹证据的情境化子声明,估计每个终端声明的置信度,最后将这些聚合为整体置信估计。在三个智能体基准、三个骨干模型和三个智能体框架上,CRG相比口头化、基于采样和白盒替代基线,产生了更好的校准置信度和更强的风险感知决策。我们进一步发现,仅校准误差可能具有误导性:一个白盒替代基线看似校准良好,却提供接近随机的判别能力。消融实验将CRG的改进归因于声明级置信估计和聚合,而非仅图构建。最后,CRG暴露了每个置信估计所依据的声明和轨迹证据,使其能在决策时被审计。
英文摘要
When using an LLM agent in a consequential domain, making an informed decision about whether to trust its output or intervene requires calibrated confidence in the agent's success. Confidence estimation for agents is difficult because evidence about success is distributed across heterogeneous, interdependent steps of an agent's trajectory. Practical agentic deployments introduce further challenges: frontier LLMs often provide limited access to internal signals, agent roll-outs are costly, and training data may be unavailable or quickly become outdated. To address these challenges, we introduce Confidence Reasoning Graphs (CRGs), an inference-time framework that estimates the probability an agent accomplished its task from a single trajectory, without privileged model access or training data. Rather than compressing an execution into a single holistic judgment, a CRG begins with the claim that the agent accomplished its task, decomposes it into contextualized sub-claims grounded in trajectory evidence, estimates confidence for each terminal claim, and finally aggregates these into an overall confidence estimate. Across three agentic benchmarks, three backbone models, and three agent frameworks, CRGs yield better-calibrated confidence and stronger risk-aware decision making than verbalized, sampling-based, and white-box surrogate baselines. We further find that calibration error alone can be misleading: a white-box surrogate baseline appears well calibrated while providing near-chance discrimination. Ablations attribute CRG's improvements to claim-level confidence estimation and aggregation rather than graph construction alone. Finally, a CRG exposes the claims and trajectory evidence underlying each confidence estimate, enabling it to be audited at decision time.
Comments34 pages, 6 figures, 11 tables