arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

规则止步之处,法官登场:多智能体系统安全中的判断边界测量

Where Rules End and Judges Begin: Measuring the Judgment Boundary in Multi-Agent Systems Security

Shaswata Mitra, Raj Patel, Subash Neupane, Sudip Mittal, Md Rayhanur Rahman, Shahram Rahimi

arXiv 2610.07657首次发表:更新:

发表机构

The University of Alabama(阿拉巴马大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出DEFER1防御框架,通过28项确定性检查与四法官评审级联,将多智能体系统攻击成功率从约30%降至3%,并揭示了规则与判断在安全防御中的边界与局限。

AI 中文摘要

基于大语言模型的多智能体系统(MAS)会调用工具、共享记忆并委派任务,因此常常遭遇对抗性内容。当前的MAS防御措施通常在隔离环境中进行评估,一次只关注一种攻击类型,这可能导致代价高昂且难以审计的结果。本研究将防御措施组织为五项原则,并将其实现为DEFER1(确定性优先执行与剩余判断),该系统包含28项检查的级联,能拦截其可拦截的内容,并将剩余部分交由四位法官组成的评审小组处理。在跨四个领域的独立测试中,攻击成功率从约30.0%降至约3.0%,其中78%的被拦截攻击由确定性检查处理。在安全运营领域,仅有四分之一的提案到达法官手中,这表明规则为违反明确政策的攻击提供了安全保障,而法官则处理那些仅歪曲意图的攻击。两种系统均存在弱点,例如风险评分审批门会不准确地批准大多数攻击提案,却拒绝少数合法提案,这凸显了准确评估威胁的挑战。

英文摘要

LLM-based multi-agent systems (MAS) engage tools, share memory, and delegate tasks, often encountering adversarial content. Current defenses for MAS are typically evaluated in isolation, focusing on one attack type at a time, which can lead to costly and hard-to-audit outcomes. This study organizes defenses into five principles, implementing them as DEFER1 (DEterministic-First Enforcement with Residual judgment), which includes a cascade of 28 checks that blocks what it can and refers the rest to a panel of four judges. In independent testing across four domains, attack success rates drop from about 30.0% to approximately 3.0%, with 78% of blocked attacks handled by deterministic checks. Only a quarter of proposals reach the judges in the security-operations domain, illustrating that the rules provide security for attacks violating clear policies, while judges manage those that only misrepresent intent. Both systems have weaknesses, such as a risk-score approval gate that inaccurately approves most attack proposals but few legitimate ones, highlighting the challenges in assessing threats accurately.

Comments26 pages, 20 figures, 24 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑