arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLM智能体系统中潜在入侵的组合式威胁分析:66号指令场景

Compositional Threat Analysis of Latent Compromise in LLM Agent Systems: The Order 66 Scenario

Satoshi Matsuoka

arXiv 2608.08131首次发表:更新:

发表机构

RIKEN Center for Computational Science(理化学研究所计算科学中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文以66号指令为场景,提出组合模型分析LLM智能体系统的潜在入侵威胁,分离三类传播路径,指出现有防御无法封闭所有路径,最强防御为能力调解等四项措施,暂未发现完整场景的公开观测。

AI 中文摘要

在虚构的66号指令中,灾难并非仅由强大指令引发:受信任群体被预先设定条件,简短指令激活隐藏状态,保护性权威转而对抗系统。本文将该机制转化为工具使用型大语言模型(LLM)智能体的起源中立安全分析。典型场景包含:带有休眠破坏规则的部署构件或共享内存、后续激活该规则的邮件、文档、更新或同伴消息,以及赋予操作与恢复权限的智能体控制模块。我们提出组合模型,解释为何各组件单独不会引发灾难,但组合后可产生关联破坏行为。我们从休眠、激活、权限、可及目标及恢复失败的共同核心中,分离出三类群体传播路径:发布前预部署、发布后持久植入、同伴复制。这形成防御割集,并说明为何检查点扫描或提示过滤无法封闭所有路径。一个两类示例显示,即使两类内复制项均低于1,跨类反馈仍可维持传播;隔离与持久控制可抑制该循环。已发表研究实例化了各组成机制,而事件显示存在自主边界跨越、恶意智能体扩展、智能体辅助侦察及公开包传播,但未出现完整的休眠植入组合。我们在2026年8月5日之前审查的证据中,未发现完整66号指令图的公开观测结果。结果既非否定也非预测:该场景在所述假设下各组件可信,损害取决于智能体控制模块,最强防御为能力调解、持久状态溯源、传播隔离及受保护恢复。

英文摘要

In the fictional Order 66, catastrophe does not arise from a powerful command alone: a trusted population is preconditioned, a short directive activates the concealed condition, and protective authority turns against the system. This paper translates that mechanism into an origin-neutral security analysis of tool-using large language model (LLM) agents. A representative scenario combines a deployed artifact or shared memory bearing a dormant destructive rule, a later email, document, update, or peer message that activates it, and an agent harness granting operational and recovery authority. We introduce a compositional model explaining why no component is catastrophic alone, yet their conjunction can produce correlated destructive action. We separate three population-reach routes --- release-time pre-positioning, post-release durable seeding, and peer replication --- from a common core of dormancy, activation, authority, reachable targets, and failed recovery. This yields defensive cut sets and shows why checkpoint scanning or prompt filtering cannot close every route. A two-class example shows that cross-class feedback can sustain spread even when both within-class reproduction terms are below one; isolation and persistence controls suppress the loop. Published work instantiates constituent mechanisms, while incidents demonstrate autonomous boundary crossing, malicious agent extensions, agent-assisted reconnaissance, and public-package propagation, but not the full dormant-implant composition. We found no public observation, in evidence reviewed through 5 August 2026, traversing the complete Order 66 graph. The result is neither dismissal nor prediction: the scenario is componentwise credible under stated assumptions, damage depends on the harness, and the strongest defenses are capability mediation, durable-state provenance, propagation isolation, and protected recovery.

CommentsVersion 11.3, 33 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑