发表机构
Pennsylvania State University; Virginia Tech; Purdue University; Amazon AGI(宾夕法尼亚州立大学; 弗吉尼亚理工大学; 普渡大学; 亚马逊通用人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型智能体的步骤级置信度估计问题,提出自进化评论家框架CEB,其通过积累自身过去判断及后果证据来估计置信度,无需训练和真实步骤标签,在多个基准和主干上校准和排名表现最佳。
AI 中文摘要
大语言模型智能体在外部环境中行动有误,单个错误步骤可能浪费交互预算或引发不可逆副作用。可靠部署需步骤级置信度估计,即行动执行前对每个提议行动有成效的校准概率。现有置信度估计器只对给定提示的响应评分,而智能体置信度还取决于执行后果。我们引入评论家经验库(CEB),是自进化评论家框架,LLM评论家积累过去判断及后果的证据。每次轨迹后,indsight LLM根据完整执行反馈对步骤是否有成效投票。结果伪标签填充记忆库,类似步骤重现时相关经验被检索到评论家提示中。CEB无需训练,不使用真实步骤标签。在三个智能体基准和三个评论家主干上,CEB在每个数据集-评论家组合中校准和排名最佳,相对最强无训练基线,ECE降低达54%。
英文摘要
LLM agents operate in stateful environments, where a single erroneous step can waste limited interaction budget or cause irreversible effects before task failure becomes apparent. Reliable deployment therefore requires step-level confidence estimation: estimating, before execution, the probability that a proposed action will advance the task. Existing LLM confidence estimators are typically designed for static question answering under a fixed task context and evaluation criterion. For an agent, however, its action productivity depends on an environment transition that is observed only after execution. To address this challenge, we introduce Critic Experience Bank (CEB), a training-free framework that turns feedback from completed trajectories into reusable evidence for future confidence judgments. After each trajectory, an LLM assigns hindsight productivity pseudo-labels to individual actions and stores them with the critic's original pre-execution confidence, task context, action, and observed feedback. For a new action, by retrieving related productive and unproductive experiences, CEB grounds pre-execution confidence in feedback from completed trajectories to condition a fixed LLM critic. CEB thereby adapts over a task stream without parameter updates or ground-truth step labels at deployment. Across four agent benchmarks spanning offline and live web navigation, mobile GUI and shell tasks, and three critic backbones, CEB achieves the best or tied-best ECE, Brier score, and AUC in all twelve benchmark-backbone settings under rule-based step labels, reducing ECE by up to 53.8% relative to the strongest training-free baseline. Its confidence scores also improve downstream utility in selective execution and simulated task success.
Comments20 pages, 5 figures