MADBench:多智能体辩论安全性基准
MADBench: Benchmarking the Security of Multi-Agent Debate
浏览论文内容
中文总结 AI 辅助
提出MADBench基准,系统评估多智能体辩论在多种攻击下的安全性,发现辩论可缓解问答准确性攻击但放大未授权读写,且多数串通攻击难以改变最终答案。
中文摘要 AI 辅助
多智能体辩论(MAD)通过允许多个智能体针对同一任务交换并批判彼此的答案,能够提升大语言模型(LLM)的推理能力。然而,这种使智能体能够纠正错误的交互机制,也可能传播对抗性错误,并将智能体引向错误答案。尽管已有一些工作考察了针对MAD的特定攻击类型,但在多种攻击下对MAD进行系统性评估仍然有限。一个核心问题是:辩论究竟是缓解了对抗性影响,还是放大了它?在本文中,我们提出了MADBench,一个用于评估MAD安全性的基准。我们按照MAD的工作流程,将攻击组织成分层分类体系,其中既包含已有的攻击方法,也包含针对辩论新设计的策略。我们在356个源任务和3,958个测试用例上评估了六个攻击家族,考察了它们对最终答案的影响以及对抗性影响的传播情况。结果表明,在攻击下,MAD并不一定能提升LLM的推理能力。与单智能体基线相比,MAD在问答任务中能够缓解对答案准确性的攻击,但在问答任务和工作区任务中都会放大未经授权的读取或写入。此外,即使五个智能体中有三个串通,攻击也仅能使原本在无攻击情况下回答正确的任务中28.30%的最终答案从正确变为错误,而辩论过程中最初正确的诚实智能体中仅有3.26%转为错误答案。
英文摘要
Multi-agent debate (MAD) can improve large language model (LLM) reasoning by allowing multiple agents to exchange and critique their answers to the same task. However, the interactions that enable agents to correct mistakes can also spread adversarial errors and steer the agents toward an incorrect answer. Although some efforts have been made to examine particular attack types on MAD, systematic evaluation of MAD under diverse attacks remains limited. A central question is whether debate mitigates adversarial influence or amplifies it. In this paper, we present MADBench, a benchmark for evaluating the security of MAD. We organize attacks into a layered taxonomy following the MAD workflow, incorporating both established attacks and new strategies tailored to debate. We evaluate six attack families over 356 source tasks and 3,958 test cases, examining their effects on the final answer and the propagation of adversarial influence. Our results show that, under attacks, MAD does not necessarily improve LLM reasoning. Compared with a single-agent baseline, MAD can mitigate attacks on answer accuracy in question-answering tasks while amplifying unauthorized reads or writes in both question-answering and workspace tasks. Moreover, even when three out of five agents collude, the attack changes the final answer from correct to wrong on only 28.30\% of tasks answered correctly without attack, while only 3.26\% of initially correct honest agents switch to wrong answers during debate.
发表机构
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。