发表机构
Boston Children’s Hospital, Harvard Medical School; Harvard School of Engineering and Applied Sciences, Harvard College(波士顿儿童医院,哈佛医学院; 哈佛工程与应用科学学院,哈佛学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过23个模型和5500多次实验证明,多智能体辩论的收益并非源于认知多样性,而是集成采样效应,在匹配预算下辩论不优于自洽采样,且角色提示存在“角色税”成本。
AI 中文摘要
多智能体辩论(MAD)据称能比单模型推理提升推理能力和事实性,但先前的工作将智能体视为对称的同伴,未阐明驱动其收益的因素。我们在问题仍可测量的场景下检验假设:智能体间的认知多样性是驱动因素,即具有基准性能提升空间的小型开放权重模型。跨越来自11个厂商家族的23个模型、5个任务以及5500多次辩论和对照运行,我们沿三个轴——角色设定、采样温度和模型身份——变化多样性,并将每个辩论配置与生成预算匹配的多数投票对照配对。该假设在每个轴上均被拒绝。辩论优于单智能体推理(在任务有提升空间处提升3至7个百分点),但在匹配预算条件下,它持平甚至输给自洽采样,同时消耗1.6倍的墙钟时间和3.4倍的令牌成本。角色提示降低了准确性,且对每个模型完整组合角色空间进行的剂量反应实验表明,代价是角色税而非多样性税:冗余角色伤害最大,而最大多样性团队能恢复部分损失。此外,混合模型团队输给其自身名单上的多数投票,准确性跟随成员能力而非异质性,且辩论的几乎所有收益来自第一轮答案交换。我们进一步识别出一个普遍的测量隐患:辩论记录会静默溢出服务上下文窗口,仅修正此问题就使我们的辩论对采样比较从-1.8个百分点变为持平。我们的结果将已报告的MAD收益重新解释为集成采样效应,并为未来辩论机制提供了必须超越的预算匹配、污染检查的基线标准。
英文摘要
Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity among agents is the driver, in the setting where the question is still measurable: small open-weight models with benchmark headroom. Across 23 models from eleven vendor families, five tasks, and 5,500+ debate and control runs, we vary diversity along three axes - personas, sampling temperature, and model identity - pairing every debate configuration with a generation-budget-matched majority-vote control. The hypothesis is rejected on every axis. Debate beats single-agent inference (3--7 points where tasks have headroom) but at matched budget conditions it ties or even loses to self-consistency sampling at 1.6$\times$ the wall-clock and 3.4$\times$ the token cost. Persona prompting reduces accuracy and a dose-response experiment over each model's full combinatorial persona space shows the cost is a persona tax, not a diversity tax: redundant personas hurt most, while maximally-diverse teams recover part of the loss. Furthermore, mixed-model teams lose to majority votes over their own rosters, with accuracy tracking member capability rather than heterogeneity, and nearly all of debate's benefit comes from the first exchange of answers. We further identify a pervasive measurement hazard in which debate transcripts silently overflow serving context windows, whose correction alone moves our debate-versus-sampling comparison from $-1.8$ points to parity. Our results recast reported MAD gains as an ensemble-sampling effect and provide the budget-matched, contamination-checked baseline bar that future debate mechanisms should be required to clear.
CommentsSubmitted to ICLR 2027