arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26212cs.SEcs.AI

多智能体辩论策略:综述、分类与挑战

Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges

  • Universitat Politècnica de Catalunya(加泰罗尼亚理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Quim Motger, Marc Oriol, Jordi Marco, Xavier Franch

AI总结:

本文通过对141项多智能体辩论研究的系统综述,提出三维分类法,揭示领域默认采用狭窄设计模式的问题,将分类法定位为研究图谱与基准框架,未来拟形式化为可执行规范以优化辩论配置。

AI中文摘要:

多智能体辩论(Multi-Agent Debate, MAD)是提升基于大语言模型(Large Language Model, LLM)的智能体系统准确性与鲁棒性的极具前景的范式,它使多个智能体能够交换论点、相互批判输出并迭代收敛至解决方案。然而,该领域的研究仍呈碎片化状态,存在术语不一致的问题,且缺乏对MAD设计维度的严谨综合。本文开展系统文献综述,对141项MAD相关基础研究进行梳理,推导得出涵盖辩论参与者、构建交换过程的交互机制以及决定辩论结果的共识协议的三维分类法,并辅以形式化符号以呈现MAD配置。分析发现,该领域已默认采用一种狭窄的设计模式——静态全连接拓扑、逐字交换、短期记忆及投票决议策略,这种模式是因惯例而非系统比较被采用,而有前景的替代方案仍处于边缘状态。由于任何MAD设置均涉及约十几个相互作用的设计决策,当这些决策未被明确时,跨研究比较不可靠。本文将该分类法定位为研究领域的描述性图谱、受控基准测试的框架,以及潜在的机器可读MAD规范架构。未来工作中,本文提出将其形式化为可执行规范,以实现成本感知型基准测试与辩论配置的自动调优。

英文摘要:

Multi-Agent Debate (MAD) is a promising paradigm for improving the accuracy and robustness of Large Language Model (LLM)-based agentic systems. It enables multiple agents to exchange arguments, critique each other's outputs, and iteratively converge towards a solution. However, research remains fragmented, with inconsistent terminology and no rigorous synthesis of MAD design dimensions. We present a systematic literature review characterizing 141 primary studies on MAD. We derive a three-dimensional taxonomy covering debate participants, the interaction mechanisms structuring the exchange, and the agreement protocols governing debate resolution, supported by formal notations to render MAD configurations. Our analysis reveals that the field has implicitly converged on a narrow design pattern - static, fully connected topologies, verbatim exchange, short-term memory and voting resolution strategies - adopted by convention rather than systematic comparison, while promising alternatives remain marginal. Because any MAD setting reflects roughly a dozen interacting design decisions, cross-study comparison is unreliable when these are left implicit. We position the taxonomy as a descriptive map of the research landscape, a framework for controlled benchmarking, and potentially as a schema for machine-readable MAD specifications. As future work, we propose formalizing it into an executable specification, enabling cost-aware benchmarking and automated tuning of debate configurations.

补充信息

↑