arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MIRROR:用于LLM多智能体通信的多路径法定人数完整性

MIRROR: Multipath Quorum Integrity for LLM Multi-Agent Communication

Ryuichi Yamafuji Lun, Jingzhen Wang, Shreyas Kolte, Ruiteng Li

arXiv 2610.02349首次发表:更新:

发表机构

University of Southern California(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MIRROR通过多路径法定人数完整性机制,在低于妥协阈值时实现零攻击成功率,同时保持低成本,为LLM多智能体通信提供了一种有效的完整性保护方案。

AI 中文摘要

智能体间通信是大型语言模型多智能体系统(LLM-MAS)的核心,但它引入了一个未被充分探索的漏洞:中间人攻击(AiTM)可以在不破坏智能体本身的情况下篡改传输中的消息。先前的研究报告在结构化任务上的攻击成功率(ASR)接近100%。现有的防御依赖于语义验证,这需要额外的推理并可能阻止良性输出,或者依赖于传输层加密,这在中间人合法终止TLS时无济于事。我们提出了MIRROR,一种通信层完整性原语,它将单个规范化载荷复制到k条逻辑路由上,并且仅当严格多数的路由报告相同的摘要时才接受消息。MIRROR使用无密钥哈希,因此本身不验证任何内容,因为活跃的路径上的对手总是可以重新计算其修改后的载荷的摘要。所有完整性都源于诚实路由构成多数的假设。摘要仅用于使见证路由保持恒定大小,并在第二原像抵抗下将恢复的载荷绑定到法定人数一致的值。我们在路由妥协界限alpha < 0.5下给出保证,并将其扩展到相关路由,其中关键的量是最大共享故障组的大小,而不是路由数量。我们进一步表明,可用性和完整性在同一阈值下降:低于alpha = 0.5时,法定人数拒绝和消息丢弃的对手无法阻止诚实流量。在MMLU、HumanEval和MBPP上,跨两个框架和四种通信拓扑,以及在MetaGPT部署中针对生产API,MIRROR在1倍LLM令牌成本下将ASR降低到阈值以下的0%。在同一部署中,LLM-as-a-Judge的成本为35倍,并在拓扑扫描中阻止了高达44.2%的良性输出。

英文摘要

Inter-agent communication is central to Large Language Model Multi-Agent Systems (LLM-MAS), but it introduces an underexplored vulnerability: Agent-in-the-Middle (AiTM) attacks that manipulate messages in transit without compromising the agents themselves. Prior work reports Attack Success Rates (ASR) approaching 100% on structured tasks. Existing defenses rely on semantic validation, which requires additional inference and can block benign outputs, or on transport-layer encryption, which does not help when an intermediary legitimately terminates TLS. We present MIRROR, a communication-layer integrity primitive that replicates a single canonicalized payload across k logical routes and accepts a message only when a strict majority of routes report the same digest. MIRROR uses unkeyed hashing and so authenticates nothing on its own, since an active on-path adversary can always recompute a digest over a payload it has modified. All integrity derives from the assumption that honest routes form a majority. The digest serves only to make witness routes constant-size and to bind the recovered payload to the quorum-agreed value under second-preimage resistance. We give the guarantee under a route-compromise bound alpha < 0.5, and extend it to correlated routes, where the quantity that matters is the size of the largest shared-failure group and not the route count. We further show that availability and integrity degrade at the same threshold: below alpha = 0.5, quorum-denial and message-dropping adversaries cannot block honest traffic. Across MMLU, HumanEval, and MBPP on two frameworks and four communication topologies, and in a MetaGPT deployment against a production API, MIRROR reduces ASR to 0% below the threshold at 1x LLM token cost. LLM-as-a-Judge costs 35x in the same deployment, and blocks up to 44.2% of benign outputs in the topology sweep.

CommentsAccepted at the NeurIPS 2026 Workshop on Foundations of Language Model Security (FLMSec). 12 pages, 4 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑