arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体交流前:多智能体系统中的事前失败风险推断

Before Agents Speak: Pre-hoc Failure Risk Inference in Multi-Agent Systems

Shi Lin, Chenpei Wang, Peng Qian, Dezhang Kong, Minghao Li, Yufeng Li, Xun Wang

arXiv 2607.26836首次发表:更新:

AI 中文总结

本研究针对多智能体系统中幻觉级联失败的问题,提出事前风险推断框架HalluProp,通过建模内在幻觉风险与传播机制实现早期诊断,性能优于事后方法,可提升系统可靠性。

AI 中文摘要

基于大语言模型(LLM)的多智能体系统(MAS)在协作推理与决策方面展现出卓越能力,但智能体间的互联通信引入了新的系统性风险:局部幻觉会沿智能体通信链传播,经交互放大后最终引发级联失败。现有对策主要遵循事后范式,仅在不安全行为出现后才识别失败,此时有害影响可能已在智能体网络中扩散。为解决该问题,本研究探索了一种互补的事前方法,提出HalluProp——一种感知传播的幻觉推断框架,用于在智能体间交互前估计单个智能体的失败情况及涌现的系统级幻觉风险。首先,通过识别智能体角色与任务查询间的细粒度语义错位,对内在幻觉风险进行建模;接着,通过对语义影响与通信拓扑同时建模,刻画智能体间的风险传播;最后,通过可微分的Noisy-OR推断机制整合上述两种风险,得出系统性诊断结果。大量实验表明,HalluProp能准确定位故障智能体,平均AUROC达84.6%,诊断耗时在亚秒级,比事后方法快65倍以上。通过在上游筛选实现早期干预,HalluProp有效补充了事后方法,凸显了事前风险推断对构建更可靠多智能体系统的潜力。

英文摘要

LLM-based multi-agent systems (MAS) have exhibited remarkable capabilities in collaborative reasoning and decision-making, yet their interconnected communications introduce new systemic risk: localized hallucinations can propagate along agent communication chain, amplify through interactions, and ultimately trigger cascading failures. Existing countermeasures predominantly follow a post-hoc paradigm, identifying failures only after unsafe behaviors emerge, by which time harmful effects may have already spread throughout the agent network. To tackle this problem, we investigate a complementary pre-hoc approach and propose HalluProp, a Propagation-aware Hallucination inference framework that estimates individual agent failures and emergent system-level hallucination risks before inter-agent interaction. First, we model intrinsic hallucination risks by identifying fine-grained semantic misalignment between agent roles and task queries. We then characterize inter-agent risk propagation by modeling both semantic influence and communication topology. Finally, we integrate these two risks via a differentiable Noisy-OR inference mechanism to derive a systemic diagnosis. Extensive experiments show that HalluProp accurately localizes faulty agents, achieving an average AUROC of 84.6%, while enabling sub-second diagnosis with over $65\times$ speedup over post-hoc methods. By facilitating early intervention through upstream screening, HalluProp effectively complements post-hoc methods, highlighting the potential of pre-hoc risk inference for building more reliable multi-agent systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑