arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22130cs.MAcs.CL

PropUQ-MAS:面向大语言模型多智能体系统的传播感知不确定性量化

PropUQ-MAS: Propagation-Aware Uncertainty Quantification for LLM Multi-Agent Systems

Yaokun Liu, Yifan Liu, Daniel Yue Zhang, Ruichen Yao, Zelin Li, Dong Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有不确定性量化方法无法捕捉大语言模型多智能体系统中不确定性传播的问题,提出PropUQ-MAS框架,实验显示其可显著提升该系统的不确定性量化性能。

中文摘要 AI 辅助

基于大语言模型(LLM)的多智能体系统(MAS)通过角色专业化智能体间的协作完成复杂任务,但智能体间的依赖关系会带来超出单个智能体故障的可靠性风险,例如中间消息的错误会被下游智能体继承并放大。现有不确定性量化(UQ)方法主要针对孤立响应或单智能体推理,无法捕捉MAS中的不确定性传播。为此,我们提出PropUQ-MAS,这是一种感知错误传播的UQ框架,将MAS执行表示为通信结构图,通过结合局部不确定性与上游消息继承的不确定性来估计每一步的可靠性。大量实验表明,PropUQ-MAS可持续提升MAS中的UQ性能,AUROC平均相对提升6.10%,PRR平均相对提升47.58%。

英文摘要

LLM-based multi-agent systems (MAS) solve complex tasks through communication among role-specialized agents. However, inter-agent dependencies introduce reliability risks beyond isolated agent failures. For instance, errors in intermediate messages could be inherited and amplified by downstream agents. Existing uncertainty quantification (UQ) methods mainly target isolated responses or single-agent reasoning, and therefore fail to capture uncertainty propagation in MAS. To this end, we propose PropUQ-MAS, an error propagation-aware UQ framework that represents MAS execution as a communication-structured graph and estimates each step's reliability by combining local uncertainty with uncertainty inherited from upstream messages. Extensive experiments demonstrate that PropUQ-MAS consistently improves UQ in MAS, with average relative gains of +6.10% in AUROC and +47.58% in PRR.

发表机构

  • Scale AI
  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑