arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MasDrift:跨多智能体架构的授权保留基准测试

MasDrift: Benchmarking Authorization Preservation Across Multi-Agent Architectures

Zhuoning Xu, Xiucheng Zhang, Hanjun Luo, Yingbin Jin, Yinpeng Dong, Hanan Salam

arXiv 2608.07556首次发表:更新:

发表机构

New York University; New York University Abu Dhabi; The Hong Kong Polytechnic University; Tsinghua University(纽约大学; 纽约大学阿布扎比分校; 香港理工大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MasDrift是含600项任务的多智能体授权保留基准,对比不同架构的任务完成与授权表现,发现集中式架构完成率更高但未授权操作更多,两种防御方法各有优劣,揭示了多智能体设计的集中化权衡。

AI 中文摘要

多智能体系统(MAS)将长周期任务分解给监督者和子智能体,但委派的目标不一定保留原始授权边界。现有安全基准主要研究对抗性破坏,而关于约束漂移的工作缺乏受控架构级评估。我们引入MasDrift,这是一个涵盖8个领域的600项良性生产力任务的基准,每项任务将所需工作与保留操作配对。MasDrift对比单智能体、集中式和分布式协调,同时改变层级深度和对等宽度,测量任务完成率和授权保留率。在通用多智能体条件下,集中式架构的任务完成率为93.9--98.6%,而对等网络为85.7--87.0%;未授权操作发生在2.7--19.8%的任务中,而对等网络为0.6--0.8%,该差距随层级深度增大而扩大。我们进一步对比两种防御方法,二者的授权证据位置不同:一种将每个待处理调用重新锚定到原始用户请求,另一种沿委派链传递衰减策略。重新锚定在我们评估的所有模型配置中均减少了未授权操作,但会使合并完成率降低1.6个百分点;链传播则会阻断所需工作,最高损失36.3个百分点。一项异构案例研究证实,该失败源于协调而非模型强度。MasDrift揭示了集中化权衡,并使授权保留成为MAS设计的可测量属性。

英文摘要

Multi-agent systems (MAS) decompose long-horizon tasks across supervisors and subagents, but delegated goals do not necessarily carry their original authorization boundaries. Existing safety benchmarks mainly study adversarial compromise, while work on constraint drift lacks controlled architecture-level evaluation. We introduce MasDrift, a benchmark of 600 benign productivity tasks across eight domains. Each task pairs required work with reserved actions. MasDrift compares single-agent, centralized, and decentralized coordination while varying hierarchy depth and peer width, measuring task completion and authorization preservation. Across generic multi-agent conditions, centralized hierarchies achieve 93.9--98.6% task completion versus 85.7--87.0% for peer networks, while unauthorized actions occur in 2.7--19.8% of tasks versus 0.6--0.8%, a gap that widens with hierarchy depth. We further compare two defenses that differ in where authorization evidence resides. One re-anchors every pending call to the original user request. The other carries an attenuated policy along the delegation chain. Re-anchoring reduces unauthorized actions in every model configuration we evaluate, at a cost of 1.6 points of pooled completion. Chain propagation blocks required work instead, forfeiting up to 36.3 points. A heterogeneous case study confirms that the failure follows from coordination rather than model strength. MasDrift exposes a centralization tradeoff and makes authorization preservation a measurable property of MAS design.

Commentspreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑