发表机构
IBM T. J. Watson Research Center(IBM托马斯·J·沃森研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一个简洁模型,基于LLM智能体的四种行为(保留异议、内化答案、重新考虑、修正)解释讨论何时提升准确性,并指出保留率低于临界值时讨论收益最大,且不保留异议或关闭推理可增加收益。
AI 中文摘要
多智能体系统(由大语言模型构成)在多数投票的基础上增加了讨论环节,因此预期其能力更强。然而,关于讨论是提高了准确性还是导致了错误的共识,实证报告结果不一。在此,我们引入一个简洁模型,该模型基于在LLM智能体中反复观察到的四种行为来解释讨论何时提高准确性以及何时以错误的共识告终:(1)保留异议(不表达不同意见),(2)内化已陈述的答案,(3)在看到异议后重新考虑,以及(4)向正确答案修正。该模型表明,只有当保留率$c$低于临界率$c^* = \gamma/(\gamma + a)$(由净修正率$\gamma$和内化率$a$决定)时,讨论才能推翻错误的初始多数。我们使用贝叶斯方法从对话日志中估计这些速率,并将LLM团队相对于$c^*$进行定位。正如模型所预测的,随着保留率的上升,讨论的收益在各类LLM以及隐藏配置文件基准(HiddenBench)和MedEInst上均减少。指示智能体不要保留异议会增加这种收益。关闭推理也会增加收益,因为推理提高了内化率$a$,并阻止智能体重新考虑少数派答案。这些发现调和了相互矛盾的报告,并确定了讨论优于多数投票的条件。
英文摘要
Multi-agent systems of LLMs add discussion to majority voting and are therefore expected to be more capable. However, empirical reports conflict on whether discussion improves accuracy or leads to an incorrect consensus. Here, we introduce a parsimonious model that explains when discussion improves accuracy and when it ends in an incorrect consensus, built from four behaviors repeatedly observed in LLM agents: (1) withholding dissent, (2) internalizing a stated answer, (3) reconsidering after seeing dissent, and (4) correcting toward the correct answer. The model shows that discussion can overturn an incorrect initial majority only when the withholding rate $c$ is below a critical rate $c^* = γ/(γ+ a)$, set by the net correction rate $γ$ and the internalization rate $a$. We estimate these rates from conversation logs with a Bayesian method and place LLM teams relative to $c^*$. As the model predicts, the gain from discussion shrinks as withholding rises, across LLMs and on a hidden profile benchmark, HiddenBench, and MedEInst. Instructing agents not to withhold dissent increases this gain. Turning reasoning off also increases the gain, because reasoning raises the internalization rate $a$ and keeps agents from reconsidering a minority answer. These findings reconcile the conflicting reports and identify when discussion outperforms majority voting.