发表机构
Illinois Institute of Technology(伊利诺伊理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究多智能体系统中分布式后门问题及局部监测器漏洞,提出可观测性边界概念,通过实验证明局部监测器在局部无害时会失效,还展示了特定监测器和门的效果,指出找到能暴露负载的表示形式是未解决问题。
AI 中文摘要
随着多智能体、使用工具的大语言模型系统的部署,一个常见的安全措施是运行时监测器,它独立检查每条消息、工具调用或步骤。我们发现这个安全措施存在一个根本性漏洞。分布式后门会将有害负载分散到多个智能体上,使得每个局部检查都通过,但组合后的对象却是攻击。监测器可能在每一步都正确,但仍会遗漏攻击。问题不在于分散本身,分散的片段仍可能泄露可疑令牌或溯源边。难题在于局部无害性,即没有片段携带危害,剩余部分看似普通的良性流量。我们将此形式化为一个可观测性边界:监测器只能捕捉其视图中能与良性流量区分开的内容。我们证明,一旦片段在监测视图中看似良性,该视图上的任何检测器都无法捕捉到它们,无论其多么强大。在一个受控测试平台、一个外部基准测试以及端到端的智能体运行中,当局部证据消失时,局部监测器会失去信号,只有当监测器看到组合后的对象时信号才会恢复。仅在良性流量上训练的监测器能在保留的编码中恢复攻击的代码结构(平均AUROC为0.874)。给定编码族的解码视图门能阻止每一次测试攻击。但仅靠看到更多是不够的:全跟踪监测器和解码器仍会失败,除非它们达到能暴露负载的表示形式。当危害具有组合性时,局部安全并非全局安全,而尚未解决的问题是找到那种表示形式。
英文摘要
As multi-agent, tool-using LLM systems are deployed, a common safety net is a runtime monitor that checks each message, tool call, or step on its own. We show this net has a fundamental hole. A distributed backdoor splits a harmful payload across agents, so every local check passes while the assembled object is the attack. The monitor can be right on every step and still miss the attack. The problem is not splitting itself: split fragments can still leak suspicious tokens or provenance edges. The hard case is \emph{local benignness}. No fragment carries the harm, and what is left looks like ordinary benign traffic. We formalize this as an \emph{observability boundary}: a monitor catches only what its view can tell apart from benign traffic. We prove that once the fragments look benign in the monitored view, no detector on that view can catch them, however strong it is. Across a controlled testbed, an external benchmark, and end-to-end agent runs, local monitors lose the signal exactly as local evidence disappears, and it returns only when the monitor sees the assembled object. A monitor trained only on benign traffic recovers the attack's code structure across held-out encodings (0.874 mean AUROC). A decoded-view gate, given the encoding family, blocks every tested attack. But seeing more is not enough: full-trace monitors and decoders still fail unless they reach the representation where the payload is exposed. Local safety is not global safety when harm is compositional, and the open problem is finding that representation.