AI 中文总结
本研究探讨单轮UQ方法能否迁移至LLM智能体的交互式轨迹场景,评估三类UQ方法在多轮工具使用数据集上的表现,发现黑盒自一致性通常最优,提示需在轨迹层面重新验证单轮UQ方法。
AI 中文摘要
语言模型的不确定性量化(UQ)方法通常在单轮输出上进行评估,此时不确定性被附加到单个生成的答案中。然而,对于大语言模型(LLM)智能体而言,观测单元是交互式轨迹,模型可提出澄清问题、调用工具、更新状态并做出中间决策,这些决策的错误会传播到最终结果。本研究探讨三类常见的单轮UQ方法是否可迁移至该场景。在来自BFCL-v4和τ²-bench的5个LLM及4个多轮工具使用数据集上,我们评估了基于动作token概率的白盒评分器、基于重采样轨迹的黑盒一致性评分器,以及基于模型对轨迹自我评估的反射式评分器。研究发现,迁移通常有用但效果不均:token概率评分对跨轮聚合器的选择高度敏感;反射式评分器在多数评估场景中提供最强的低成本基线;黑盒自一致性常是最强的UQ类别,其变体中轨迹等价性和动作集一致性通常排名最高。这些结果表明,为单轮生成开发的UQ方法应在轨迹层面重新验证,需密切关注一致性测量、聚合器选择及计算预算。
英文摘要
Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying questions, call tools, update state, and make intermediate decisions whose errors propagate to the final outcome. We study whether three common families of single-turn UQ methods transfer to this setting. Across five LLMs and four multi-turn tool-use datasets from BFCL-v4 and $τ^2$-bench, we evaluate white-box scorers based on action-token probabilities, black-box consistency scorers based on resampled trajectories, and reflexive scorers based on model self-assessment of the trajectory. We find that transfer is often useful but uneven. Token-probability scores are highly sensitive to the choice of aggregator used across turns, reflexive scores provide the strongest low-cost baseline in most evaluated settings, and black-box self-consistency is often the strongest UQ family, with trajectory-equivalence and action-set consistency typically ranking highest among its variants. These results suggest that UQ methods developed for single generations should be revalidated at the trajectory level, with careful attention to the consistency measurement, aggregator choice, and computational budget.