ORBIT:多智能体安全与安保评估框架
ORBIT: A Framework for Multi-Agent Safety and Security Evaluations
浏览论文内容
中文总结 AI 辅助
ORBIT是一个基于Inspect构建的可配置多智能体安全评估框架,支持多种威胁与防御策略,发现防御措施在不同威胁间迁移性不足,并揭示安全-性能权衡及架构与防御有效性的交互。
中文摘要 AI 辅助
多智能体大语言模型系统正越来越多地被部署用于复杂、长期任务,或作为智能体在真实环境中交互的自然结果而出现。然而,它们也带来了重大的安全与安保风险:使任务泛化成为可能的灵活协议同样暴露了新的威胁,从级联提示注入到智能体间共谋。防御这些威胁的进展因缺乏共享的经验基础设施而放缓,这迫使每种新防御都需要定制环境开发,并使标准化比较变得不可能。现有评估针对孤立的威胁模型或单智能体场景,但没有一种评估能在现实的多智能体环境中联合变化攻击、防御和架构。为弥补这一空白,我们引入了ORBIT,一个用于经验性多智能体安全与安保研究的可配置评估框架,构建于英国AISI的Inspect之上。ORBIT允许研究人员配置通信拓扑、内存、调度和智能体角色。它支持四种威胁类型和四种防御策略,以及非对抗性故障,并提供一个涵盖五个场景族(包括浏览器使用、计算机使用、智能体编码、客户服务和协作分配)的基准套件。我们的核心发现是防御在威胁间的可迁移性存在差距:将受损智能体在多问题编码中的攻击成功率降低60个百分点的逐动作防御,对共谋智能体没有可测量的保护作用,并且我们测试的防御中没有一个能在所有测试的攻击上泛化。我们进一步展示了安全-性能权衡以及架构与防御有效性之间的交互。我们在该https URL上开源了ORBIT。
英文摘要
Multi-agent LLM systems are increasingly deployed for complex, long-horizon tasks or emerge as a natural consequence of agents interacting in the wild. Yet they give rise to significant safety and security risks: the flexible protocols that enable task generalization also expose novel threats, from cascading prompt injection to inter-agent collusion. Progress in defending against these threats has been slowed by a lack of shared empirical infrastructure, which forces bespoke environment development for every new defense and makes standardized comparison impossible. Existing evaluations address isolated threat models or single-agent settings, but none jointly vary attack, defense, and architecture across realistic multi-agent environments. To address this gap, we introduce ORBIT, a configurable evaluation framework for empirical multi-agent safety and security research, built on UK AISI's Inspect. ORBIT lets researchers configure communication topologies, memory, scheduling, and agent roles. It supports four threat types and four defense strategies, as well as non-adversarial failures, with a benchmark suite spanning five scenario families covering browser use, computer use, agentic coding, customer service, and cooperative allocation. Our central finding is a gap in defense transferability across threats: per-action defenses that cut a compromised agent's attack success by 60 points on multi-issue coding give no measurable protection against colluding agents, and none of the defenses we tested generalized over all attacks tested. We further demonstrate security-performance tradeoffs and interactions between architecture and defense effectiveness. We make ORBIT available open-source at https://github.com/wlanderson0/orbit.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
- MATS Research(MATS 研究)
- University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。