MiniRep:面向多智能体辩论的鲁棒基于声誉的聚合方法
MiniRep: Robust Reputation-Based Aggregation for Multi-Agent Debate
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
MiniRep提出一种面向多智能体辩论的鲁棒声誉聚合方法,结合当前行为与历史声誉,并抑制相似响应群体主导,在MATH任务上优于传统聚合和声誉方法。
AI中文摘要:
由大型语言模型(LLM)驱动的自主智能体正在迅速发展成为一个开放的智能体生态系统。为了支持可信赖的协作,行业倡议越来越多地根据过去的行为评估智能体的声誉,并提供性能排行榜。然而,基于过去表现的声誉可能无法可靠地预测智能体在新任务上的行为,尤其是当恶意智能体能够调整自身行为并在协作过程中影响其他智能体时。我们研究了多智能体辩论(MAD)中的声誉问题,其中多个智能体回答同一查询,通过辩论改进各自的答案,并将这些答案聚合为最终输出。我们提出了MiniRep,一种在存在恶意智能体的MAD场景下基于声誉的聚合系统。为了将我们的威胁模型建立在已有研究的基础上,我们构建了一个攻击分类体系,借鉴了声誉系统攻击和软件测试变异算子,涵盖了声誉的战略性利用和对智能体提案的微妙篡改。在该分类体系的指导下,MiniRep基于智能体在当前任务上的行为及其随时间积累的声誉来评估智能体,同时防止具有高度相似响应的智能体群体主导最终决策。我们跨多种任务、LLM智能体组成、篡改位置以及源自我们分类体系的攻击类型对MiniRep进行了评估。实验结果表明,在MATH数据集上,无论是否受到攻击,MiniRep都优于传统的MAD聚合方法和传统的基于声誉的方法。此外,在MATH数据集上异构10智能体设置下,MiniRep在所有28种攻击条件下均优于所有基线方法。
英文摘要:
Autonomous agents powered by large language models (LLMs) are rapidly evolving into an open agentic ecosystem. To support trustworthy collaboration, industry initiatives increasingly assess agent reputation from past behavior and provide performance leaderboards. However, reputation derived from past performance may not reliably predict an agent's behavior on new tasks, particularly when malicious agents can adapt their behavior and influence other agents during collaboration. We study reputation in multi-agent debate (MAD), where multiple agents answer the same query, debate to improve their answers, and aggregate them into a final output. We present MiniRep, a reputation-based aggregation system for MAD under malicious agents. To ground our threat model in established research, we construct an attack taxonomy drawing on reputation-system attacks and software-testing mutation operators, covering strategic exploitation of reputation and subtle corruption of agent proposals. Guided by this taxonomy, MiniRep evaluates agents based on both their behavior on the current task and their reputation over time, while preventing groups of agents with highly similar responses from dominating the final decision. We assess MiniRep across diverse tasks, LLM-agent compositions, corruption placements, and attack types drawn from our taxonomy. Our experimental results show that, MiniRep outperforms both conventional MAD aggregation and conventional reputation-based approaches on MATH no matter being attacked or not. Also, under a heterogeneous 10-agent setting on MATH, MiniRep outperforms all baselines in all 28 attack conditions.