从序列到结构:面向大语言模型智能体的关系型不确定性传播
From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents
浏览论文内容
中文总结 AI 辅助
本研究针对现有LLM智能体不确定性量化方法忽略长程依赖的问题,提出RUPA框架,通过建模轨迹图的关系依赖传播不确定性,在多个基准上实现更优性能。
中文摘要 AI 辅助
可靠的不确定性量化(UQ)对于将大语言模型(LLM)智能体部署到复杂交互环境中至关重要。现有的UQ方法大多依赖局部信号,如token概率、预测熵或每步置信度,因此忽略了错误在执行轨迹中累积的长程依赖关系,导致无法识别出原因源自最终答案前若干推理或交互步骤的智能体故障。我们提出RUPA(Relational Uncertainty Propagation for Agents,面向智能体的关系型不确定性传播),这是一种面向LLM智能体的轨迹级UQ框架。RUPA将执行历史表示为有向轨迹图,其中推理状态、工具交互和环境反馈是节点,由时间和语义依赖边连接;随后在该图上传播不确定性,以捕捉执行风险如何在交互步骤间累积和转移。传播信号与轨迹级行为特征及目标对齐信息相结合,生成整个智能体轨迹的置信度估计。我们在代表性智能体基准(包括τ-2、Terminal-Bench-2和GAIA)上,使用覆盖多个模型家族的6种开源LLM对RUPA进行评估。实验结果表明,RUPA通过提供更准确的不确定性估计、实现更早的故障检测以及在各类智能体任务中改进不确定性引导的智能体执行,始终优于现有UQ方法。这些结果证明,显式建模关系依赖对长程LLM智能体的可靠UQ至关重要,为可信智能体执行提供了实用基础。
英文摘要
Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments. Existing UQ methods largely rely on local signals, such as token probabilities, predictive entropy, or per-step confidence, and therefore overlook the long-range dependencies through which errors accumulate across an execution trajectory. As a result, they may fail to identify agent failures whose causes originate several reasoning or interaction steps before the final answer. We propose RUPA (Relational Uncertainty Propagation for Agents), a trajectory-level UQ framework for LLM agents. RUPA represents an execution history as a directed trajectory graph in which reasoning states, tool interactions, and environment feedback are nodes connected by temporal and semantic dependency edges. It then propagates uncertainty over this graph to capture how execution risk accumulates and transfers across interaction steps. The propagated signal is combined with trajectory-level behavioral features and goal-alignment information to produce a confidence estimate for the full agent trajectory. We evaluate RUPA on representative agent benchmarks, including $τ$-2, Terminal-Bench-2, and GAIA, using 6 open-source LLMs spanning multiple model families. Experimental results show that RUPA consistently outperforms existing UQ methods by providing more accurate uncertainty estimates, enabling earlier failure detection, and improving uncertainty-guided agent execution across diverse agent tasks. These results demonstrate that explicitly modeling relational dependency is crucial to reliable UQ for long-horizon LLM agents, providing a practical foundation for trustworthy agent execution.
发表机构
- Institute of Software, Chinese Academy of Sciences(中国科学院软件研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。