arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21259cs.AI

CogGym:迈向人类与机器认知的大规模比较评估

CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition

  • Massachusetts Institute of Technology(麻省理工学院)
  • Harvard University(哈佛大学)
  • Cornell University(康奈尔大学)
  • Johns Hopkins University(约翰斯·霍普金斯大学)
  • Helmholtz Munich(亥姆霍兹慕尼黑中心)
  • Massachusetts General Hospital(马萨诸塞总医院)
  • University of Cambridge(剑桥大学)
  • Princeton University(普林斯顿大学)
  • Santa Fe Institute(圣塔菲研究所)
  • Dartmouth College(达特茅斯学院)
  • EPFL(洛桑联邦理工学院)
  • University of Chicago(芝加哥大学)
  • University of Tübingen(蒂宾根大学)
  • CHI-FRO
  • Stanford University(斯坦福大学)
  • University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
  • University of British Columbia(不列颠哥伦比亚大学)
  • Vector Institute(向量研究所)
  • Canada CIFAR AI Chair(加拿大CIFAR人工智能讲席)
  • Yale University(耶鲁大学)
  • University of California, Berkeley(加利福尼亚大学伯克利分校)
  • University of Washington(华盛顿大学)
  • New York University(纽约大学)
  • McGill University(麦吉尔大学)
  • Mila–Quebec AI Institute(米拉-魁北克人工智能研究所)
  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)
  • Prior Computers
  • Purdue University(普渡大学)
  • National University of Singapore(新加坡国立大学)
  • A*STAR Institute of Advanced Intelligence and Computing(新加坡科技研究局先进智能与计算研究所)
  • The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

Lance Ying, Jinzhou Wu, Yingshan Susan Wang, Shivam Aarya, Luca M. Schulze Buschoff, Harry Chen, Katherine M. Collins, Andrea de Varda, Shuhao Fu, Sean Dae Houl… 展开作者

Lance Ying, Jinzhou Wu, Yingshan Susan Wang, Shivam Aarya, Luca M. Schulze Buschoff, Harry Chen, Katherine M. Collins, Andrea de Varda, Shuhao Fu, Sean Dae Houlihan, Akshay K. Jagadish, Guangyuan Jiang, Samuel Kiegeland, Tetsu Kurumisawa, Rongzhi Liu, Ryan Liu, Ningshan Ma, Kathryn McGregor, Younes Strittmatter, Polina Tsvilodub, Jacob Hoover Vigly, Sarah Wu, Enjie Xu, Yiling Yun, Kelsey Allen, Tyler Brooke-Wilson, Brian Christian, Evelina Fedorenko, Michael C. Frank, Michael Franke, Tao Gao, Samuel J. Gershman, Robert D. Hawkins, Jennifer Hu, Julian Jara-Ettinger, Max Kleiman-Weiner, Sydney Levine, Tal Linzen, Hongjing Lu, Timothy O'Donnell, Desmond C. Ong, Steven T. Piantadosi, Rebecca Saxe, Eric Schulz, Tianmin Shu, Felix A. Sosa, Ilia Sucholutsky, Tan Zhi-Xuan, Tomer Ullman, Fei Xu, Ilker Yildirim, Jian-Qiao Zhu, Thomas L. Griffiths, Tobias Gerstenberg, Kevin Smith, Joshua B. Tenenbaum

AI总结:

CogGym提出一个基于认知科学的可扩展统一框架,通过标准化258个实验并评估50个大语言模型,发现模型与人类判断的拟合度随规模提升但远低于人类自身信度,且进步慢于形式推理基准。

AI中文摘要:

理解和建模人类智能是人工智能(AI)与认知科学共同追求的并行目标。随着AI系统能力日益增强,模型响应在哪些方面与人类响应相似,又在哪些方面系统性偏离?人类能够执行和思考的任务范围之广、种类之多,对人类与模型的可扩展且严格的比较构成了挑战。我们提出CogGym,一个基于认知科学的可扩展、统一框架,用于在匹配的实验试次上系统比较模型与人类行为。CogGym采用半自动、人在回路的流水线,将多样化的实验范式标准化为任务无关的实验标记语言(EML),从而实现在规模上的可复现和忠实的比较。在初始版本中,我们从100篇论文中整理并标准化了258个认知实验,这些实验聚焦于人类常识推理,并评估了50个大语言模型与人类响应的匹配度。我们发现一个清晰的缩放趋势:更大、更新的AI模型能更好地复现人类判断。然而,AI模型在这些常见推理任务上的提升速度明显慢于在数学和编程等正式推理基准上的进步,且模型与人类的拟合度仍远低于人类自身的分半信度(文本为$R^2 = 0.93$,图像为$0.95$,视频为$0.92$),最佳模型在文本实验上达到$R^2 = 0.59$,图像为$0.58$,视频为$0.43$。我们期望CogGym能提供一个活的评估框架,持续纳入新的认知科学实验,以刻画模型行为在何处与人类行为相似、何处系统性偏离,以及这些模式如何随模型和实验的演进而变化。

英文摘要:

Understanding and modeling human intelligence are parallel goals shared by artificial intelligence (AI) and cognitive science. As AI systems grow increasingly capable, in what ways do model responses resemble human responses, and where do they systematically diverge? The sheer breadth and diversity of the tasks humans can perform and think about pose a challenge for scalable and rigorous comparison between humans and models. We introduce CogGym, a scalable, unified framework grounded in cognitive science for systematically comparing model and human behavior on matched experimental trials. CogGym uses a semi-automated, human-in-the-loop pipeline to standardize diverse experimental paradigms into a task-agnostic Experiment Markup Language (EML), enabling reproducible and faithful comparison at scale. For initial release, we curate and standardize 258 cognitive experiments from 100 papers that focuses on human commonsense reasoning, and evaluate 50 large language models against human responses. We find a clear scaling trend where larger and more recent AI models better reproduce human judgments. Yet AI models' improvement on such common reasoning tasks is considerably slower than the gains observed on formal-reasoning benchmarks like math and coding, and model--human fit remains well below human splithalf reliability ($R^2 = 0.93$ on text, $0.95$ on image, and $0.92$ on video) with the best models achieving $R^2 = 0.59$ on text, $0.58$ on image, and $0.43$ on video experiments. We intend for CogGym to provide a living evaluation framework that continually incorporates new cognitive science experiments to characterize where model behavior resembles human behavior, where it systematically diverges, and how those patterns change as models and experiments evolve.

补充信息

↑