arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

知形察因:多智能体大语言模型故障的拓扑条件诊断

Know the Shape, Find the Fault: Topology-Conditioned Diagnosis of Multi-Agent LLM Failures

Xinwen Liu, Zhuocheng Pan, Isabella Zhu, Jawei Zhang, Xudong Liu, Tianyu Wo

arXiv 2610.10126首次发表:更新:

发表机构

Beihang University(北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出MAScope两阶段框架,利用通信拓扑结构线索诊断多智能体LLM故障,通过轨迹结构提取器恢复拓扑并条件化分类,显著提升诊断性能并降低部署成本。

AI 中文摘要

多智能体大语言模型(LLM)系统通过智能体之间的信息交换来协调任务执行。当协调失败时,执行轨迹中的相似症状可能反映出信息传递、使用或验证方面的不同问题。通信拓扑刻画了智能体如何交换信息,并为区分协调失败模式提供了结构线索。利用这些线索进行诊断,需要建立拓扑与故障模式之间的关系,并从缺乏显式拓扑标签的执行轨迹中恢复相关结构。我们分析了通信拓扑与故障模式之间的关系,并引入了MAScope,一个用于拓扑条件诊断的两阶段框架。其轨迹结构提取器(TSE)通过将交互图基于消息证据进行 grounding,从异构执行轨迹中恢复通信拓扑。拓扑条件法官(TC-Judge)则利用轨迹、预测的拓扑、从独立标记轨迹中估计的经验故障先验,以及拓扑特定故障模式的简短描述来对故障进行分类。在固定的编排结构下,恢复的拓扑可跨多次执行重用。实验结果显示,通信拓扑与故障类型之间存在统计学显著关联,χ² = 409.9,p = 1.2 × 10⁻⁷⁰。在851条MAST-clean轨迹上,真实拓扑上下文将gpt-mini的Macro-F1从0.173提升至0.350。使用预测拓扑时,该流程达到0.346,接近仅使用轨迹的gpt-5.4基线0.372。对于固定编排结构下的1000条轨迹,预计流程成本(包括一次拓扑提取)约为重复gpt-5.4诊断成本的6%。这些结果表明,拓扑条件上下文改善了故障诊断,并支持更低成本的部署。

英文摘要

Multi-agent LLM systems coordinate task execution through exchanges of information among agents. When coordination breaks down, similar symptoms in execution traces can reflect different problems in how information is passed, used, or verified. Communication topology captures how agents exchange information and provides structural cues for distinguishing coordination failure modes. Using these cues for diagnosis requires establishing how topology relates to failure patterns and recovering the relevant structure from execution traces that lack explicit topology labels. We analyze the relationship between communication topology and failure patterns and introduce MAScope, a two-stage framework for topology-conditioned diagnosis. Its Trace Structural Extractor TSE recovers communication topology from heterogeneous execution traces by grounding an interaction graph in message evidence. The Topology-Conditioned Judge TC-Judge then classifies failures using the trace, predicted topology, an empirical failure prior estimated from separate labeled traces, and a short description of topology-specific failure patterns. Under a fixed orchestration structure, the recovered topology can be reused across executions. Experimental results show a statistically significant association between communication topology and failure type, with $χ^2 = 409.9$ and $p = 1.2 \times 10^{-70}$. On the \num{851} MAST-clean traces, ground-truth topology context raises gpt-mini's Macro-F1 from $0.173$ to $0.350$. With predicted topology, the pipeline achieves $0.346$, approaching the trace-only gpt-5.4 baseline of $0.372$. For \num{1000} traces under a fixed orchestration structure, the projected pipeline cost, including one topology extraction, is approximately $6\%$ of repeated gpt-5.4 diagnosis cost. These results show that topology-conditioned context improves failure diagnosis and supports lower-cost deployment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑