发表机构
University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出匹配轨迹重放协议,对比原始置信度与事后等渗校准在多跳问答系统中的表现,发现校准可提升承诺回答准确率但影响整体准确率,且无法估计检索预期收益。
AI 中文摘要
交互式语言模型智能体利用置信信号决定是立即回答、检索额外证据(来自记忆或外部知识)还是弃权(不执行)。然而,置信度通常是单独评估的,未测量其触发的动作在轨迹层面的后果。我们提出匹配轨迹重放,这是一种用于比较置信度到动作映射的受控协议。该协议保持候选答案状态、证据点、预算和动作成本固定。我们用它比较原始口头置信度与事后等渗校准在多跳问答系统中的表现,该系统使用Mistral、GPT和Qwen模型,在HotpotQA和MuSiQue数据集上运行。在相同的数值承诺阈值下,校准会改变智能体最终承诺回答的问题。在所有6组模型-数据集组合中,它将承诺回答的准确率最多提高41个百分点,但会降低覆盖范围并增加检索使用。在HotpotQA上的整体准确率最多提高15个百分点,但在MuSiQue上最多下降17个百分点。这些影响反映了向更具选择性、更低风险的操作点转变,而非答案或置信度排名的改进。在检索深度1和2时,检索前拟合的校准图改善了保留的校准,但在深度3时,对于所有三个模型,其表现均差于原始置信度。额外证据平均有帮助,但这种总体效应无法确定置信度是否能识别出哪些单个样本将受益于另一次检索。综上,这些结果表明,校准可以使承诺风险可解释,但无法估计另一次检索的预期收益,因此检索需要单独的信息价值或效用估计。评估应报告保留的校准、风险-覆盖范围和检索成本。
英文摘要
Interactive language-model agents use confidence signals to decide whether to answer immediately, retrieve additional evidence (from memory or external knowledge), or defer. Yet confidence is usually evaluated in isolation, without measuring the trajectory-level consequences of the actions it triggers. We propose matched trajectory replay, a controlled protocol for comparing confidence-to-action mappings. The protocol holds candidate answer states, evidence points, budgets, and action costs fixed. We use it to compare raw verbalized confidence with post-hoc isotonic calibration in a multi-hop question-answering system using Mistral, GPT, and Qwen models on HotpotQA and MuSiQue datasets. At the same numerical commitment threshold, calibration changes which questions agents ultimately commit to answering. Across all six model-dataset pairs, it increases accuracy among committed answers by up to 41 percentage points. However, it can reduce coverage and increase retrieval use. Overall accuracy improves by up to 15 percentage points on HotpotQA but falls by up to 17 percentage points on MuSiQue. These effects reflect a shift to a more selective, lower-risk operating point, not improved answers or confidence ranking. A calibration map fitted before retrieval improves held-out calibration through retrieval depths one and two, but is worse than raw confidence at depth three for all three models. Additional evidence helps on average, but this aggregate effect does not establish whether confidence identifies which individual episodes will benefit from another retrieval. Taken together, these results show that calibration can make commitment risk interpretable, but it does not estimate the expected benefit of another retrieval. Retrieval therefore requires a separate value-of-information or utility estimate. Evaluations should report held-out calibration, risk-coverage, and retrieval cost.
Comments5 Tables, 5 Figures