arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2510.07505cs.LG

PEAR:规划-执行代理鲁棒性基准

PEAR: Planner-Executor Agent Robustness Benchmark

  • Michigan State University(密歇根州立大学)
  • Purdue University(普渡大学)
  • University of Texas at Dallas(德克萨斯大学达拉斯分校)

机构由 AI 辅助整理,请以论文原文为准。

Shen Dong, Mingxuan Zhang, Pengfei He, Li Ma, Bhavani Thuraisingham, Hui Liu, Yue Xing

更新

AI总结:

PEAR基准通过评估规划-执行多智能体系统的鲁棒性,揭示了规划器弱点对任务性能的影响及鲁棒性与性能的权衡。

AI中文摘要:

基于大型语言模型(LLM)的多智能体系统(MAS)已经发展成为一种强大的范式,用于解决复杂、多步骤任务,跨越各种领域。然而,尽管其表现令人印象深刻,MAS仍然容易受到对抗性操纵的影响。现有研究通常只考察孤立的攻击面或特定场景,缺乏对MAS脆弱性的全面理解。为填补这一空白,我们引入PEAR,一个系统评估规划-执行MAS实用性和脆弱性的基准。虽然兼容各种MAS架构,我们的基准专注于规划-执行结构,这是一种实用且广泛采用的设计。通过广泛的实验,我们发现:(1)弱化的规划器对整体干净任务性能的影响比弱化的执行器更严重;(2)虽然记忆模块对规划器至关重要,但为执行器添加记忆模块不影响干净任务性能;(3)任务性能和鲁棒性之间存在权衡;(4)针对规划器的攻击特别有效于误导系统。这些发现为增强MAS的鲁棒性提供了可操作的见解,并为多智能体设置中的原则性防御奠定了基础。

英文摘要:

Large Language Model (LLM)-based Multi-Agent Systems (MAS) have emerged as a powerful paradigm for tackling complex, multi-step tasks across diverse domains. However, despite their impressive capabilities, MAS remain susceptible to adversarial manipulation. Existing studies typically examine isolated attack surfaces or specific scenarios, leaving a lack of holistic understanding of MAS vulnerabilities. To bridge this gap, we introduce PEAR, a benchmark for systematically evaluating both the utility and vulnerability of planner-executor MAS. While compatible with various MAS architectures, our benchmark focuses on the planner-executor structure, which is a practical and widely adopted design. Through extensive experiments, we find that (1) a weak planner degrades overall clean task performance more severely than a weak executor; (2) while a memory module is essential for the planner, having a memory module for the executor does not impact the clean task performance; (3) there exists a trade-off between task performance and robustness; and (4) attacks targeting the planner are particularly effective at misleading the system. These findings offer actionable insights for enhancing the robustness of MAS and lay the groundwork for principled defenses in multi-agent settings.

补充信息

↑