arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多智能体大语言模型系统中分布式后门的早期检测:一项特征研究

Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study

Diego Fernandez Arias, Dev Prashant Mistry, Ren Wang, Yibo Hu

arXiv 2607.24893首次发表:更新:

发表机构

Illinois Institute of Technology(伊利诺伊理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多智能体大语言模型系统中分布式后门早期检测,构建分层多智能体系统实例,通过记录片段注入及载荷组装执行时间来检测。前缀检测器能较早标记多数成功攻击,部分检测器依赖表面线索,微调模型可减少线索去除后的损失。

AI 中文摘要

多智能体大语言模型系统可能会受到一种单个智能体都不会完整持有载荷的攻击:一个中毒工具在其观测中隐藏加密片段,将它们分散到多个智能体中,并且在运行后由外部步骤重新组装并执行。逐步骤的安全检查单独判断每个动作可能无法识别完整的分布式载荷。我们研究在运行仍在进行时,这种攻击能多早被检测到,以及一旦其最明显的线索被去除,它能多稳健地被捕获。我们在分层多智能体系统上构建了一个工作实例,在五个语言模型和两个任务域的良性和受攻击条件下运行它,并记录每个片段何时被注入以及载荷何时被组装和执行。检测是与组装的竞赛。在第一个片段被注入之前,受攻击和良性运行是无法区分的;一旦注入开始,一个前缀检测器能标记出99.3%的成功攻击,中位数为还剩五步,安全运行误报率为10.3%。由于组装仅在运行后发生,这些警报能及时到达以中止几乎每一次成功攻击。然后我们测量该警告有多少依赖于攻击的可去除表面线索而非其分布式结构。通用的零样本和行为训练检测器几乎不提供任何警告;起作用的检测器部分依赖于可去除表面线索,主要是密文的长度和熵,一旦从载荷中去除熵线索以及从检测器中去除长度特征,检测会延迟且跨域传递性差,不过一个微调模型能恢复一些损失。

英文摘要

Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: several poisoned tools each hide one encrypted fragment, spreading them across several agents, and an external step reassembles and executes them after the run. Per-step safety checks that judge each action in isolation may fail to recognize the complete distributed payload. We investigate how early such an attack can be detected while the run is still unfolding, and how robustly it can be caught once its most obvious cues are stripped away. We build a working instance on a hierarchical multi-agent system, run it under benign and attacked conditions across five language models and two tool environments, and record when each fragment is injected and when the payload is assembled and executed. Almost no run is flagged before its first fragment is injected; once injection begins, a prefix detector flags $99.5\%$ of successful attacks with a median of twelve steps remaining and an out-of-fold false-alarm rate of about $1\%$ on uninjected runs, higher on runs that carried fragments but never assembled. Because assembly occurs only after the run, these alarms could enable an abort before assembly on nearly every successful attack. We also test a detector that sees only the current observation, to ask whether earlier steps are needed. When the current observation carries an obvious clue, one observation is often enough; when that clue is removed, using earlier steps can help. Generic zero-shot and behavior-trained detectors fail to separate unsafe from safe runs; the detectors that do work lean in part on removable surface cues, chiefly the ciphertext's length and entropy. Take those cues away and the detector fires later, and carrying it from one tool environment to the other becomes much harder. A model fine-tuned on the raw text recovers part of the loss.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑