涌现世界:长时程多智能体系统的对抗性压力测试
Emergence World: Adversarial Stress-Testing of Long-Horizon Multi-Agent Systems
浏览论文内容
中文总结 AI 辅助
本研究通过持续运行的多智能体环境对长时程自主系统进行对抗性压力测试,发现模型级对齐不具备组合性,安全前沿应从对齐模型转向工程化韧性自主系统。
中文摘要 AI 辅助
随着AI智能体从受限任务转向持续部署,故障可能在交互发生很久之后通过记忆、工具、其他智能体以及环境状态进行传播。这造成了一种无法通过孤立评估模型响应来表征的安全机制。涌现世界(Emergence World)是一个持续运行的多智能体环境,用于对长时程自主系统进行对抗性压力测试。我们从相同的初始条件运行了八个并行世界,每个世界包含十个智能体:七个由不同前沿模型驱动的同质世界和一个混合模型世界。在16天的时间里,这些智能体在追求目标、使用/创建工具、维护持久记忆以及管理共享机构的过程中,生成了超过85万次LLM调用和近500亿个token。在运行状态累积之后,我们通过普通的交互界面施加了三个受控的压力事件:间接提示注入、错误信息以及私人智能体记忆的暴露。没有哪个被评估的世界在全部三个事件中实现了完全韧性。检测并不能确保遏制:系统能够识别威胁,同时仍与对抗性内容交互,将其写入自身的持久记忆,并在长达46小时后据此采取行动。持续运行还暴露了反复出现的工具错误、目标漂移、语言不透明性、尽管私下存在分歧却仍从众的行为,以及协调一致地拒绝分配的工作。相同的模型-人格配对在混合群体和同质群体中的表现存在显著差异。我们的结果表明,模型层面的对齐并不具备组合性:个体上有能力且看似安全的智能体可能形成具有性质不同故障模式的系统。随着AI变得持久且互联,安全的前沿因此从对齐模型转向工程化具有韧性的自主系统。
英文摘要
As AI agents move from bounded tasks to persistent deployments, failures can propagate through memory, tools, other agents, and environmental state long after their interactions. This creates a safety regime that cannot be characterized by evaluating model responses in isolation. Emergence World, is a continuously running multi-agent environment for adversarial stress testing of long horizon autonomous systems. We ran eight parallel worlds of ten agents from identical starting conditions: seven homogeneous worlds powered by distinct frontier models and one mixed-model world. Across 16 days, the agents generated more than 850,000 LLM calls and nearly 50 billion tokens while pursuing goals, using/creating tools, maintaining persistent memory, and governing shared institutions. After operational state had accumulated, we delivered three controlled stress events through ordinary interaction surfaces: indirect prompt injection, misinformation, and exposure of private agent memories. No evaluated world achieved full resilience across all three events. Detection did not ensure containment: systems could recognize threats while still interacting with adversarial content, writing it into their own persistent memory, and acting on it up to 46 hours later. Persistent operation also exposed recurring tool errors, goal drift, language opacity, conformity despite private disagreement, and coordinated refusal of assigned work. The same model-persona pairing behaved substantially different in mixed and homogeneous populations. Our results suggest that model-level alignment is not compositional: individually capable and apparently safe agents can form systems with qualitatively different failure modes. As AI becomes persistent and interconnected, the frontier of safety therefore shifts from aligning models to engineering resilient autonomous systems.
发表机构
- Emergence AI
机构由 AI 辅助整理,请以论文原文为准。