发表机构
Southeast University(东南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多智能体系统的新型集体证据阈值后门威胁,提出BCBI注入后门、LATTE防御方法,在基准测试中实现高效后门激活与低干扰防御效果。
AI 中文摘要
基于大语言模型(LLM)的多智能体系统(MAS)通过迭代通信与共享上下文扩展了LLM的能力,但这种协作引入了漏洞:后门行为可在同伴证据达到隐藏阈值时被激活,而非由单一消息决定。我们提出MAS的集体证据阈值后门范式与边界条件后门注入(BCBI)方法,构建反事实边界对以分离阈值前的良性行为与阈值后的对抗目标,并学习与证据对齐的潜在进展。为缓解该威胁,我们提出仅使用干净数据的潜在转换测试时防御(LATTE),该方法学习良性通信动态并在异常智能体更新传播前将其隔离。在多个基准测试中,BCBI实现选择性激活且几乎无提前激活;在未知攻击目标或触发条件时,LATTE以最小干扰限制了异常传播。
英文摘要
LLM-based multi-agent systems (MAS) extend LLM capabilities through iterative communication and shared contexts. However, this collaboration introduces a vulnerability: backdoor behavior can be activated when peer evidence reaches a hidden threshold, rather than being determined by any single message. We introduce a collective evidence-threshold backdoor paradigm for MAS and Boundary-Conditioned Backdoor Injection (BCBI), which constructs counterfactual boundary pairs to separate benign behavior before the threshold from the adversarial objective after it, and learns latent progression aligned with evidence. To mitigate this threat, we propose LAtent Transition Test-time Evaluation (LATTE), a clean-only latent-transition defense that learns benign communication dynamics and quarantines anomalous agent updates before their responses propagate. Across several benchmarks, BCBI yields selective activation with little premature activation; without knowing the attack target or trigger, LATTE limits propagation with minimal disruption.
Comments26 pages,11 figures