发表机构
University of Zurich; ETH Zurich(苏黎世大学; 苏黎世联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究通过审议推理指数(DRI)评估1980次五智能体LLM运行,发现LLM群体讨论程序性质量与人类相当但主体间一致性提升小、视角多样性仅为人类的1/3,反转人类收敛模式,仅可作为人类推理的辅助工具而非自主审议智能体。
AI 中文摘要
大型语言模型(LLMs)正越来越多地被部署到需要对复杂、涉及价值的问题进行集体推理的场景中。人们对这些部署的信心很大程度上基于可验证任务(数学、编码、协调博弈)的基准,但许多这类应用涉及不存在客观正确答案的问题,决策质量取决于整合多元视角以找到相互可接受的解决方案。我们认为,无法从可验证任务基准中完全推断出LLM在这类问题上的推理能力,且对LLM讨论的程序性评估(尊重、论证、参与)存在系统性不足。我们采用了审议推理指数(DRI)——一个在政治学中开发并经公民大会验证的指标——作为评估多元、不可验证问题上可靠群体水平推理的工具。综合最近关于1980次五智能体LLM运行的证据,这些运行涉及11种前沿模型配置下的12个公民大会主题,我们发现LLM群体产生的讨论程序性质量与人类审议相当,但主体间一致性的提升很小、取决于主题,且集中在可处理而非伦理争议的问题上。LLM群体的视角多样性约为人类大会的三分之一,且反转了人类的收敛模式:人类审议会随着多元视角整合而减少分散性,而LLM审议会增加分散性。通过角色提示(persona prompting)实现的工程多样性无法恢复人类的动态,反而会反转审议推理的更新组成部分。我们的结论是约束性而非禁止性的:LLMs可作为支持人类对多元问题推理的工具,但现有证据不允许将它们视为自主审议智能体。
英文摘要
LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests largely on benchmarks for verifiable tasks (mathematics, coding, coordination games), yet many of these applications concern problems where no objectively correct answer exists and where decision quality instead depends on integrating pluralistic perspectives to find mutually acceptable solutions. We argue that LLM reasoning capacity on this class of problems cannot be fully inferred from verifiable-task benchmarks, and that procedural evaluations of LLM discourse (respectfulness, justification, engagement) are systematically insufficient. We apply the Deliberative Reason Index (DRI), a measure developed in political science and validated across citizen assemblies, as a tool for evaluating reliable group-level reasoning on pluralistic, non-verifiable problems. Synthesizing recent evidence across 1,980 five-agent LLM runs on 12 citizen-assembly topics across 11 frontier model configurations, we find that LLM groups produce discourse with procedural quality comparable to human deliberation, while gains in intersubjective consistency are small, topic-dependent, and concentrated on tractable rather than ethically contested questions. LLM groups exhibit roughly one-third the perspective diversity of human assemblies and reverse the human convergence pattern: human deliberation decreases dispersion as diverse views synthesise, whereas LLM deliberation increases it. Engineering diversity through persona prompting does not restore the human dynamic but inverts which component of deliberative reasoning is updated. Our conclusion is constraining rather than prohibitive: LLMs can function as tools supporting human reasoning on pluralistic problems, but current evidence does not license treating them as autonomous deliberative agents.
Comments10 pages, archival publication at AIES 2026