AI 中文总结
针对多智能体辩论中LLMs易盲目从众的问题,提出DEAR框架,通过动态调控辩论关系缓解该问题,实验显示其性能更优且token消耗显著降低。
AI 中文摘要
多智能体辩论(Multi-Agent Debate, MAD)通过多轮交互提升大语言模型(Large Language Models, LLMs)的推理性能,但MAD中的LLMs极易受盲目从众影响。现有基于置信度或困惑度的个体评估方法无法反映推理的正确性,甚至可能加剧盲目从众。为解决该问题,我们将视角从个体评估转向群体交互,定义LLMs间的相互引用为“辩论关系”,并认识到调控这些关系是缓解盲目从众的关键。本文提出一种从群体视角动态调控辩论关系的新框架——DEAR(Dynamically Regulating Debate Relationships)。首先,DEAR将共识与分歧量化为“群体证据”以捕捉辩论状态;随后,DEAR通过三个阶段运行:1)“什么”:感知群体咨询倾向与不确定性;2)“谁”:引入Selection RL-Agent以动态选择参考对等体;3)“如何”:采用Behavior RL-Agent自适应调整生成行为。值得注意的是,我们将两个RL-Agents的执行建模为序贯决策过程,通过多智能体强化学习联合优化。大量实验表明,DEAR在实现更优性能的同时,显著降低了token消耗。
英文摘要
Multi-Agent Debate (MAD) improves the reasoning performance of Large Language Models (LLMs) through multi-round interaction. However, LLMs in MAD are highly susceptible to blind conformity. Existing individual evaluation methods, typically based on confidence or perplexity, fail to reflect the correctness of reasoning and may even exacerbate blind conformity. To address this, we shift the perspective from individual evaluation to group interaction. We define mutual referencing among LLMs as \textbf{Debate Relationships} and recognize that regulating these relationships is the key to mitigating blind conformity. In this paper, we propose a novel framework for \textbf{D}ynamically r\textbf{E}gulating deb\textbf{A}te \textbf{R}elationships (DEAR) from the group perspective. At first, DEAR quantifies consensus and divergence as \textit{group evidence} to capture the debate state. Then, DEAR operates through three stages: 1) What: perceiving group consultation tendency and uncertainty; 2) Who: introducing a Selection RL-Agent to dynamically select reference peers; and 3) How: adopting a Behavior RL-Agent to adaptively adjust generation behaviors. Notably, we formulate the execution of the two RL-Agents as a sequential decision-making process, jointly optimizing via multi-agent reinforcement learning. Extensive experiments demonstrate that DEAR achieves superior performance while significantly reducing token consumption.