arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语言模型群体在推理任务上回放人类讨论时高估共识

Language-model groups overstate consensus when replaying human deliberation on a reasoning task

Tengfei Shao

arXiv 2609.20543首次发表:更新:

发表机构

Waseda University(早稻田大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过回放人类Wason推理讨论,发现语言模型智能体群体在多种评分标准下均显著高估共识,且模拟共识不反映集体准确性,为评估模拟群体估计提供了评分明确的依据。

AI 中文摘要

完全一致率常被视为集体认知的指标,然而其取决于参与和最终状态的操作化定义。我们使用匹配的大型语言模型(LLM)智能体群体回放了100个留出的人类Wason群体,每个参与者的讨论前答案中植入一个信念锚定智能体,并使用相同的代码对智能体和人类进行评分。在不同的人类评分定义下,估计值介于24.0%到57.0%之间;约五分之一的参与者从未发帖,而智能体几乎总是发帖。在两次揭盲后的敏感性分析中,智能体群体仍然表现出更高的共识:基于提交的比较(n = 98)在聊天和推理模式下分别产生了34.0和43.9个百分点的差距,而参与匹配的比较(n = 45)分别产生了34.1和44.4个百分点的差距。这些互补的路径减少了不同的测量不对称性,但收敛在0.5个百分点以内。在没有提前停止以及重新参数化去除可记忆答案的情况下,差距依然存在;推理模式群体随后几乎一致同意,且大多同意错误答案。模拟共识并未追踪集体准确性,信念锚定智能体群体在此设置中是人类群体结果分布的有偏估计量。这些分析为评估模拟群体对人类审议结果的估计提供了评分明确的依据。

英文摘要

Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant's pre-discussion answer and scoring agents and people with the same code. Across human scoring definitions, estimates ranged from 24.0% to 57.0%; about one fifth of participants never posted, whereas agents almost always did. Agent groups remained more consensual in two post-unblinding sensitivity analyses: the submit-based comparison (n = 98) yielded gaps of 34.0 and 43.9 percentage points for chat and reasoning modes, and the participation-matched comparison (n = 45) yielded gaps of 34.1 and 44.4 points. These complementary routes reduced different measurement asymmetries yet converged within 0.5 percentage points. The gap persisted without early stopping and under a reparameterization removing the memorizable answer; reasoning-mode groups then agreed nearly unanimously, mostly on incorrect answers. Simulated consensus did not track collective accuracy, and belief-anchored agent groups were biased estimators of the human group-outcome distribution in this setting. These analyses provide a scoring-explicit basis for assessing simulated-group estimates of human deliberative outcomes.

Comments37 pages, 4 figures. Preregistration: https://osf.io/5jp7s . Code and data: https://doi.org/10.5281/zenodo.21318346

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑