发表机构
SND / CNRS / Sorbonne University(SND / 法国国家科学研究中心 / 索邦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究探讨大语言模型群体是否有群体智慧,通过让15个LLMs对254个二元预测市场问题进行概率估计,评估经典和学习型聚合方法,发现学习型聚合器表现优,训练截止污染影响结果,表明LLM群体有群体智慧效应,但可靠评估需无污染评估。
AI 中文摘要
群体智慧——聚合个体判断往往优于最佳个体的发现——已在人类预测者中得到广泛研究。当“群体”由大语言模型(LLMs)组成时,同样的现象是否出现是一个具有理论和实践意义的开放性问题。我们从15个LLMs中获取了关于254个二元预测市场问题的概率估计,并评估了经典和学习型聚合方法。学习型聚合器——多层感知器和逻辑回归——优于所有个体模型和经典方法。发现逻辑回归与神经网络相当,这表明学习型聚合的好处来自学习不同模型输出的线性组合而非非线性交互。符号回归应用于神经网络的学习映射,在帕累托前沿恢复了一个纯模型分歧信号作为最低复杂度的有用公式,进一步支持了这一解释。训练截止污染被证明是一个普遍存在的混淆因素:在所有模型训练截止后解决的问题的干净子集中,前沿云模型和较小本地模型之间的明显能力差距从35.8%缩小到8.9%,个体模型排名仅显示出适度的稳定性。即使在每个模型的训练截止时评估预测市场,LLMs仍然准确性低得多,表明在集体信息聚合方面存在真正差距。这些发现表明LLM群体可以表现出群体智慧效应,但无污染评估对于可靠评估至关重要。
英文摘要
The wisdom of crowds -- the finding that aggregating judgments across individuals often outperforms the best individual -- has been extensively studied with human forecasters. Whether the same phenomenon emerges when the ``crowd'' consists of large language models (LLMs) is an open question with both theoretical and practical implications. We elicited probability estimates from 15 LLMs on 254 binary prediction market questions and evaluated classical and learned aggregation methods. Learned aggregators -- a multilayer perceptron and a logistic regression -- outperformed all individual models and classical methods. The logistic regression was found to match the neural network, suggesting that the benefit of learned aggregation derives from learning a linear combination of diverse model outputs rather than from nonlinear interactions. Symbolic regression applied to the neural network's learned mapping recovered a pure model-disagreement signal as the lowest-complexity useful formula on the Pareto frontier, further supporting this interpretation. Training cutoff contamination proved a pervasive confound: the apparent capability gap between frontier cloud models and smaller local models collapsed from 35.8% to 8.9% on a clean subset of questions resolving after all models' training cutoffs, and individual model rankings showed only moderate stability. Even when the prediction market is evaluated at each model's training cutoff, LLMs remained substantially less accurate, indicating a genuine gap in collective information aggregation. These findings suggest that LLM crowds can exhibit wisdom-of-crowds effects, but that contamination-free evaluation is essential for reliable assessment.