发表机构
Johns Hopkins University; New York University(约翰斯·霍普金斯大学; 纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大语言模型智能体在知识冲突下的认知谦逊展开评估,提出ISE三维度定义认知谦逊,发现高准确率未必对应高认知谦逊,模型干预可提升认知谦逊但常牺牲任务准确率。
AI 中文摘要
当检索到的证据与智能体的先验信念相矛盾时,它会修改答案、承认不确定性,还是坚持错误结论?现有对智能体系统的评估主要聚焦于任务成功度,对智能体如何处理此类冲突的洞察有限。我们提议评估智能体的认知谦逊(epistemic humility, EH):即智能体在任务执行过程中识别、应对并传达不确定性的意愿。我们通过轨迹层面的三个行为维度来定义认知谦逊:识别(Identify)、解决(Solve)和升级(Escalate,简称ISE)。我们设置了知识冲突场景,其中包括两种情况:一是骨干语言模型的参数化知识与其遇到的证据相矛盾,二是两个上下文来源存在分歧;我们评估两种冲突设置:(1)受控冲突,(2)多步智能体执行过程中自然出现的冲突,每种设置都配有匹配的无冲突对照组。通过评估四个智能体,我们发现更高的任务准确率并不一定对应更大的认知谦逊:一些高准确率配置会在执行过程中识别冲突,但不会在错误的最终答案中传达未解决的不确定性。轨迹层面分析进一步显示,智能体经常在执行早期步骤中检测到冲突,但在后续步骤中无法维持或解决这些冲突。最后,我们表明模型层面的干预措施可以提升认知谦逊,但往往以牺牲任务准确率为代价,这表明认知谦逊是骨干模型、智能体框架和评估环境之间相互作用的产物。
英文摘要
When retrieved evidence contradicts an agent's prior beliefs, does it revise its answer, acknowledge uncertainty, or persist with an incorrect conclusion? Existing evaluations of agentic systems focus primarily on task success, offering limited insight into how agents handle such conflicts. We propose to evaluate agents on epistemic humility (EH): the agent's willingness to recognize, act on, and communicate uncertainty during task execution. We operationalize EH through three trajectory-level behavioral dimensions: Identify, Solve, and Escalate (ISE). Through knowledge conflict, situations where the backbone language model's parametric knowledge contradicts the evidence it encounters, or where two contextual sources disagree, we evaluate two conflict settings: (1) controlled conflict and (2) naturally occurring conflict during multi-step agentic execution, each paired with matched no-conflict controls. Evaluating four agents, we find that higher task accuracy does not necessarily correspond to greater epistemic humility: some high-accuracy configurations recognize conflicts during execution but do not communicate unresolved uncertainty in their incorrect final answers. Trajectory-level analysis further reveals that agents frequently detect conflicts in early steps of execution but fail to maintain or resolve them in later steps. Finally, we show that model-level interventions can improve EH, but often at the cost of task accuracy, suggesting that epistemic humility emerges from the interaction among the backbone model, agent harness, and evaluation environment.
CommentsEMNLP 2026 Camera Ready
Journal refEMNLP2026