arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SoK:当安全智能体共同失效时:多智能体大语言模型(LLM)系统的安全性

SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems

Rui Yang, Junjie Xu, Zhengyu Liu, Neil Fendley, Yang Hong, Ziyang Li, Yinzhi Cao

arXiv 2609.00595首次发表:更新:

发表机构

Johns Hopkins University; Nanyang Technological University(约翰斯·霍普金斯大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过分析197项研究系统化梳理多智能体大语言模型(MAS)安全性,提出A-I-R框架统一攻击机制,明确防御关键挑战,审计44项基准研究并指出开放问题,推动形成感知交互的MAS安全视角。

AI 中文摘要

安全智能体可能会共同失效。多智能体大语言模型(MAS)在不同主体边界间传递信息、状态、决策和权限,会产生局部检查可能遗漏的失效情况。若缺乏执行层面的视角,多智能体场景极易被误判为真正的多智能体安全效应的证据。因此,我们通过对197项研究进行以执行为中心的分析,系统化梳理MAS安全性,涵盖6种交互接口、4种 adversary 位置、7种系统级风险及8种反复出现的攻击路径。我们引入A-I-R框架,该框架按 adversary 位置、交互接口及产生的系统级风险对攻击进行组织,统一了MAS中原本零散的攻击机制。我们通过涵盖路径目标、观测、干预、信任边界和恢复的五部分合约来组织防御措施,并确定路径闭合与恢复是关键挑战。我们对44项评估与基准研究进行审计,发现在隔离交互效应、设计可比较的诊断指标、支持MAS设计间的复用及评估开放系统运行方面存在开放挑战。这些发现共同推动形成一种感知交互的MAS安全视角:端到端追踪攻击、测试防御是否闭合这些路径,并使用适当的反事实评估系统级效应。

英文摘要

Safe agents can fail together. Multi-agent LLM systems (MAS) move information, state, decisions, and authority across principal boundaries, creating failures that local checks may miss. Without an execution-level view, a multi-agent setting can easily be mistaken for evidence of a genuinely multi-agent security effect. We thus systematize MAS security through an execution-centered analysis of 197 works, covering six interaction interfaces, four adversary positions, seven system-level risks, and eight recurring attack paths. We introduce an A-I-R framework that organizes attacks by adversary position, interaction interface, and resulting system-level risk, unifying otherwise fragmented attack mechanisms across MAS. We organize defenses through a five-part contract covering path target, observation, intervention, trust boundary, and recovery, and identify path closure and recovery as key challenges. We audit 44 evaluation and benchmark works and identify open challenges in isolating interaction effects, designing comparable and diagnostic metrics, supporting reuse across MAS designs, and evaluating open-system operation. Together, these findings motivate an interaction-aware view of MAS security: trace attacks end to end, test whether defenses close those paths, and evaluate system-level effects with appropriate counterfactuals.

Comments21 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑