发表机构
INSEAD; University of North Texas; Stanford University; Tel Aviv University; University of Waterloo; Reichman University; University of Toronto; Ben-Gurion University of the Negev; Massachusetts Institute of Technology; University of Pennsylvania; Cornell University; University of California, Berkeley; New College of Florida; Harvard University; New York University; Northwestern University; London Business School; ETH Zurich; Oregon State University(欧洲工商管理学院; 北得克萨斯大学; 斯坦福大学; 特拉维夫大学; 滑铁卢大学; 莱克曼大学; 多伦多大学; 本-古里安大学; 麻省理工学院; 宾夕法尼亚大学; 康奈尔大学; 加利福尼亚大学伯克利分校; 佛罗里达新学院; 哈佛大学; 纽约大学; 西北大学; 伦敦商学院; 苏黎世联邦理工学院; 俄勒冈州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究比较25个LLM与人类科学家在社会科学理论构建中的表现,发现AI在个体任务上更优但理论复杂且预测效率低,人类理论更简洁高效且多样,两者互补,AI擅长处理复杂性,人类多样性促进集体智慧。
AI 中文摘要
我们研究了人工智能(AI)——特别是大型语言模型(LLMs)——相对于人类科学家在社会科学高级认知任务中的有效性,这些任务包括理论构建、对新实证结果的预测以及根据新证据进行理论修正。研究领域涉及关于性别和种族不平等的学术讨论。我们的发现,通过比较25个LLMs与13位资深研究人员及60位博士学者,表明在大多数现有任务中,AI在个体层面超越了大多数人类,而人类的理论更为多样,并且在聚合时预测准确性的提升更大。AI生成的理论更为详尽,涉及额外的理论路径和潜在变量,并且被不知来源的独立评估者评为比人类理论质量更高。然而,这种理论复杂性部分是装饰性的,因为它与对数据中经验模式的更准确预测无关;相比之下,人类科学家通过更简单的理论实现了更高的预测效率。AI比人类科学家更有可能修订其理论以纳入新证据;人类科学家则以一种对先前预测错误敏感的选择性方式更新其信念。我们推测,人工智能优越的处理能力使其特别适合处理需要应对复杂性的任务,但人类思想的更大多样性对于智慧人群和集体创造力至关重要。
英文摘要
We investigate the effectiveness of artificial intelligences (AI)-specifically large language models (LLMs)-relative to human scientists at high-level cognitive tasks in social science such as theory formulation, predictions of novel empirical results, and theory revision in response to new evidence. The research domain was academic discourse regarding gender and race inequality. Our findings, comparing 25 LLMs with 13 senior researchers and 60 doctoral scholars, reveal that the AIs outperformed most humans individually on most of the present tasks, while human theories were more diverse and exhibited greater gains in predictive accuracy from aggregation. AI-generated theories were more extensively elaborated, involving additional theoretical paths and latent variables, and were rated as higher quality than human theories by independent raters blinded to source. However, this theoretical complexity was in part ornamental, in that it was not associated with more accurate predictions about empirical patterns in data; in contrast, human scientists achieved greater predictive efficiency with simpler theories. The AIs were significantly more likely than human scientists to revise their theories to incorporate new evidence; human scientists updated their beliefs in a selective way that is sensitive to prior prediction errors. We speculate that the superior processing capacity of artificial intelligences makes them especially well-suited to tasks requiring grappling with complexity, but that the greater diversity of human ideas is essential to wise crowds and collective creativity.