arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36556cs.AI

MAADBench:多智能体系统异常检测的可刷新范式

MAADBench: The Refreshable Paradigm for Anomaly Detection in Multi-Agent Systems

Lei Ma, Dennis Hofmann, Haowen Xu, Joshua DeOliveira, Peter VanNostrand, Lei Cao, Elke Rundensteiner

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM多智能体系统异常检测缺乏可刷新基准的问题,提出MAADBench,结合生成任务、可刷新轨迹与自动步骤级标签,评测25种方法揭示现有方法局限。

中文摘要 AI 辅助

近期研究报道,基于大语言模型(LLM)的多智能体系统(MAS)的失败率高达41%-87%,然而据我们所知,目前尚无基准测试能够支持对其系统性的异常检测(AD)。构建MAS AD基准测试十分困难,因为随着LLM系统的发展,基准测试必须保持新鲜:任务可能泄漏到训练数据中,从而被LLM记忆;随着骨干模型的发展,轨迹和异常模式会过时;并且每次刷新时都必须可靠地提供标签。为应对这些挑战,我们提出了MAADBench(MA:多智能体;AD:异常检测),这是首个面向支撑智能体的多样化、不断演进的LLM骨干模型设计的可刷新MAS AD基准测试。MAADBench结合了(1)在约10^37任务空间上的采样耦合生成任务以缓解任务泄漏,(2)在可配置的LLM骨干模型下可刷新的轨迹生成,以及(3)自动提供零成本、确定性的步骤级标签,用于细粒度的AD评估。除了提供范式本身,我们还使用五个最先进的LLM骨干模型运行了MAADBench,并发布了包含5,200条步骤级标记轨迹的MAADBench-Full数据集。在MAADBench数据集上对25种AD方法进行基准测试,揭示了当前方法的重大局限性:它们严重依赖监督,难以应对微妙的MAS特定异常,并且跨LLM骨干模型的鲁棒性不足。这些差距指向了针对MAS特定异常检测的丰富研究议程,而MAADBench为方法开发和评估提供了一个系统且可刷新的测试平台。我们在该https URL上开源了MAADBench-Full。

英文摘要

Recent studies report that LLM-based multi-agent systems (MAS) fail at rates of 41%-87%, yet to our knowledge, no benchmark to date supports systematic anomaly detection (AD) for them. Building MAS AD benchmarks is hard because they must remain fresh as LLM systems evolve: tasks may leak into training data and thus be memorized by LLMs, traces and anomaly patterns expire as backbones evolve, and labels must be provided reliably for each refresh. To address these challenges, we present MAADBench (MA: multi-agent; AD: anomaly detection), the first refreshable MAS AD benchmark designed for diverse, evolving LLM backbones underlying the agents. MAADBench combines (1) sampled-and-coupled generative tasks over an approximately 10^37-task space to mitigate task leakage, (2) refreshable trace generation under configurable LLM backbones, and (3) automated provision of cost-free, deterministic step-level labels for fine-grained AD evaluation. Beyond offering the paradigm itself, we run MAADBench with five state-of-the-art LLM backbones and release the MAADBench-Full dataset with 5,200 step-labeled traces. Benchmarking 25 AD methods on the MAADBench dataset reveals substantial limitations in current approaches: they rely heavily on supervision, struggle with subtle MAS-specific anomalies, and lack robustness across LLM backbones. These gaps point to a rich research agenda for MAS-specific anomaly detection, with MAADBench providing a systematic and refreshable testbed for method development and evaluation. We open-source MAADBench-Full at https://huggingface.co/datasets/hww123/MAADBench-full.

发表机构

  • Worcester Polytechnic Institute(伍斯特理工学院)
  • University of Arizona(亚利桑那大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑