arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将规则归纳带回流体智力研究?人类中ARC-AGI基准的初步验证

Bringing Back Rule Induction to Fluid Intelligence Research? An Initial Validation of the ARC-AGI Benchmark in Humans

Jasmin Thelen, Oliver Wilhelm

arXiv 2607.11263首次发表:更新:

发表机构

Ulm University(乌尔姆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究围绕流体智力测量的两种观点,以含100名参与者的研究初步验证ARC-AGI基准用于人类流体智力测量的有效性,发现其与图形流体智力显著相关,为相关研究提供支持并建议嵌入人工智能基准以促进评估与合作。

AI 中文摘要

关于流体智力(gf)测量存在两种相互竞争的观点,一种认为表现主要受工作记忆容量限制,另一种认为受归纳新关系能力限制。目前前者在测量中占主导,后者多体现在定义中。ARC-AGI基准主要要求规则归纳,被提议作为人类和人工系统的gf测量方法,但尚未在人类样本中检验其心理测量特性。本研究对100名参与者调查了ARC-AGI的心理测量特征和法则网络。ARC-AGI项目汇编显示出良好心理测量特性,与图形推理测试测量的图形流体智力显著相关(\r{ho}=.63),与图形原创性关联较弱。这些发现为ARC-AGI作为人类流体智力测量方法的有效性提供了初步支持。未来研究应纳入更多规则归纳任务及多变量协变量。本研究通过研究最初为机器设计的人类任务而与众不同,建议将人工智能基准系统地嵌入人类认知能力的法则网络,以实现更系统的评估和跨学科合作。

英文摘要

Two competing perspectives on fluid intelligence (gf) measures propose that performance is primarily constrained either by working memory capacity or by the ability to induce novel relations. The first perspective is currently dominant in measurement, as evident from the use of a limited set of recurring rules, whereas the second perspective is reflected in many definitions but rarely present in measurement. The ARC-AGI benchmark predominantly requires rule induction and was proposed as a measure of gf for both humans and artificial systems. However, its psychometric properties have not yet been examined in human samples. We therefore investigated the psychometric characteristics and nomological network of ARC-AGI in a first study with 100 participants. A compilation of ARC-AGI items showed good psychometric properties and correlated substantially with figural fluid intelligence as measured by a figural reasoning test ($ρ$ = .63). Associations with figural originality were weak. These findings provide initial support for the validity of ARC-AGI as a measure of human fluid intelligence. Future research should include more rule induction tasks as well as additional multivariate covariates. This study is unusual by studying a task in humans that was initially designed for machines. We suggest systematically embedding AI benchmarks into the nomological network of human cognitive abilities to enable more systematic evaluation and interdisciplinary cooperation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑