arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

分层多智能体系统中基于语义不确定性的协调策略

Semantic Uncertainty-Guided Orchestration in Hierarchical Multi-Agent Systems

John Knowlton, Aritra Guha, Risto Miikkulainen

arXiv 2608.14707首次发表:更新:

发表机构

The University of Texas at Austin; AT&T Chief Data Office; Cognizant AI Lab(德克萨斯大学奥斯汀分校; AT&T首席数据办公室; 高知特人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出HASSUM框架,利用语义熵和密度估计不确定性实现多智能体自适应协调,在StrategyQA等基准上验证其可提升复杂推理任务的可靠性。

AI 中文摘要

随着基于大语言模型(LLM)的多智能体系统能力日益增强,在不确定性条件下协调智能体成为一项基础性挑战。现有协调策略通常依赖固定的交互模式,且往往缺乏评估中间推理步骤可靠性的机制,导致错误和幻觉在系统中传播。本文提出一种基于语义不确定性的协调方法,将HASSUM作为多智能体系统中用于不确定性感知协调的通用框架。该方法利用语义熵和语义密度估计不确定性,这两个指标在答案语义层面而非输出概率层面衡量可信度。这些信号支持自适应协调决策,包括输出验证、选择性重新提示、额外审议以及置信度感知的响应选择。由于该方法独立于任何特定智能体架构运行,因此可集成到广泛的分层及协作多智能体系统中。评估在分层智能体框架内实现该方法,并在StrategyQA、JailbreakBench和TruthfulQA基准上进行测试。在需要复杂推理且易出现歧义或幻觉的任务中,与不考虑不确定性的协调方式相比,基于不确定性的协调策略产生更可靠的结果。语义熵与语义密度结合的效果优于单独使用任一指标。对不同阈值和模型规模的消融实验表明,两者均会影响语义指标的有效性。结果表明,语义不确定性是一种实用且通用的信号,可用于提升智能体AI系统的鲁棒性和可信度。

英文摘要

As large language model (LLM)-based multi-agent systems become increasingly capable, coordinating agents under uncertainty becomes a fundamental challenge. Existing orchestration strategies typically rely on fixed interaction patterns and often lack mechanisms for assessing the reliability of intermediate reasoning steps, allowing errors and hallucinations to propagate through the system. This paper introduces a semantic-uncertainty-guided orchestration approach, HASSUM as a general framework for uncertainty-aware coordination in multi-agent systems. The method estimates uncertainty using semantic entropy and semantic density, which measure trust at the level of answer semantics rather than output probabilities. These signals enable adaptive orchestration decisions, including output verification, selective reprompting, additional deliberation, and confidence-aware response selection. Because the approach operates independently of any particular agent architecture, it can be integrated into a broad range of hierarchical and collaborative multi-agent systems. The evaluations demonstrate an implementation within a hierarchical agent framework and evaluate it on StrategyQA, JailbreakBench, and TruthfulQA benchmarks. Across tasks that require complex reasoning and are prone to ambiguity or hallucinations, uncertainty-guided orchestration yields more reliable outcomes than uncertainty-unaware coordination. Semantic entropy and semantic density in tandem outperformed either metric alone. Ablations testing different thresholds and model sizes demonstrated that both influence the effectiveness of semantic metrics. The results suggest that semantic uncertainty is a practical and general-purpose signal for improving robustness and trustworthiness in agentic AI systems.

Comments17 pages, 5 figures, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑