arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24831cs.AI

GRUET:量化智能体推理与行动过程的不确定性

GRUET: Quantifying Uncertainty of Agentic Reasoning-and-Acting Processes

Shuang Liang, Xin-Yu Hu, Shao-Qun Zhang

AI总结:

针对智能体ReAct过程中轨迹级不确定性导致的可信度问题,提出GRUET方法,通过图建模推理分支空间量化逐轮不确定性并聚合为轨迹级不确定性,在九个LLMs和五个基准上验证了有效性。

AI中文摘要:

智能体因能够在开放和动态环境中执行推理与行动(ReAct)而受到越来越多的关注。ReAct过程通常表现出多轮轨迹,其中驱动大型语言模型(LLMs)以交错方式生成推理链和特定任务的动作。然而,智能体常常面临显著的不确定性,导致相同任务产生不同的轨迹;具有较高不确定性的轨迹往往产生难以理解的行为,严重削弱了智能体的可信度。本研究推测,这种轨迹级别的不确定性通常源于LLMs引起的累积的逐轮推理不确定性;后者通常表现为一系列发散推理链及其相应动作的分支集合。基于此,我们提出了基于图的轨迹推理不确定性(GRUET)方法,用于ReAct的不确定性量化,包括逐轮推理不确定性量化和轨迹级不确定性聚合;前者通过将潜在推理分支所跨越的推理空间建模为图,然后用图的复杂度近似推理空间的复杂度,从而精确量化推理不确定性,而后者采用简单的聚合策略来量化整体轨迹的可信度。在九个LLMs和五个基准上的实证评估验证了我们提出的GRUET在选择性生成性能方面的有效性,通过AUROC、AUPRC和AUARC来衡量。

英文摘要:

Agents have attracted considerably increasing attention due to the power of executing both Reasoning and Acting (ReAct) in open and dynamic environments. The ReAct process typically exhibits a multi-turn trajectory in which one drives Large Language Models (LLMs) to generate both reasoning chains and task-specific actions in an interleaved manner. However, agents often suffer from significant uncertainty, where identical tasks yield divergent trajectories; trajectories with higher uncertainty often produce incomprehensible behaviors, severely undermining agent credibility. This work conjectures that such trajectory-level uncertainty frequently stems from cumulative turn-level reasoning uncertainty induced by LLMs; the latter often exhibits a collection of branches of divergent reasoning chains and their resulting actions. Built upon this, we present the Graph-based Reasoning UncErtainty in Trajectories (GRUET) method for the uncertainty quantification of ReAct, comprising turn-level reasoning uncertainty quantification and trajectory-level uncertainty aggregation; the former precisely quantifies reasoning uncertainty via modeling the reasoning space spanned by potential reasoning branches as a graph and then approximating the reasoning space complexity with graph complexity, while the latter employs simple aggregation strategies for quantifying the overall trajectory credibility. Empirical evaluations across nine LLMs and five benchmarks validate the effectiveness of our proposed GRUET in terms of selective generation performance, measured by AUROC, AUPRC, and AUARC.

↑