发表机构
Santa Fe Institute(圣塔菲研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对标准ARC评估仅关注单一能力的问题,提出多维基准PotARCin,从定义、分类等五维度评估抽象技能习得,发现模型性能显著下降且排名改变,并引入P-ARC测试集凸显全面评估的重要性。
AI 中文摘要
抽象与推理语料库(ARC)已成为评估AI模型通用抽象推理和流体智能的重要基准。然而,标准ARC评估仅考虑单一能力:为测试输入生成正确的输出网格。我们认为,这种狭窄的格式无法评估真正抽象技能习得所应促成的多种能力。我们引入了PotARCin,一个扩展ARC的基准,通过五个维度评估对任务潜在抽象规则的理解:定义、分类、约束生成、编辑和反演。PotARCin采用程序化方法为给定的ARC任务生成新任务实例并转换给定输入,从而实现了超越固定输入-输出对的动态生成采样。在ARC-AGI-1训练集上评估的五个最先进模型中,我们观察到标准ARC评估与PotARCin评估之间存在25-52个百分点的性能差距,并发现多维评估重新排序了标准准确率排名相同的模型。我们进一步研究了生成采样、损坏类型难度以及自一致性问题的影响,表明模型经常与自己形式化的规则相矛盾,即使它们已正确陈述了该规则。我们还引入了P-ARC,一个留出的手工构建测试集,模型在该测试集上所有五个维度的准确率仅为1-8%,这强调了更全面评估抽象推理能力的重要性。
英文摘要
The Abstraction and Reasoning Corpus (ARC) has become a prominent benchmark for evaluating general abstract reasoning and fluid intelligence in AI models. Yet standard ARC evaluation considers only a single capability: producing the correct output grid for a test input. We argue that this narrow format fails to evaluate the diversity of abilities that genuine abstract skill acquisition should enable. We introduce PotARCin, a benchmark that extends ARC by assessing understanding of a task's underlying abstract rule across five dimensions: Definition, Classification, Constrained Generation, Editing, and Inversion. PotARCin employs programmatic methods to generate new task instances and transform given inputs for a given ARC task, enabling dynamic generative sampling beyond fixed input-output pairs. Across five state-of-the-art models evaluated on the ARC-AGI-1 training set, we observe a 25-52 percentage-point performance gap between standard ARC evaluation and evaluation on PotARCin, and find that multi-dimensional evaluation reorders models that standard accuracy ranks alike. We further investigate effects of generative sampling, difficulty of corruption types, and questions of self-consistency, showing that models frequently contradict their own formalized rule even where they have stated it correctly. We also introduce P-ARC, a held-out hand-crafted test set, on which models achieve 1-8% accuracy across all five dimensions, underscoring the importance of more holistic evaluations of abstract reasoning capabilities.