arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无系统性的思维?评估推理模型在规则归纳任务上的表现

Thought without systematicity? Evaluating reasoning models on rule induction tasks

Simon Schug, Brenden M. Lake

arXiv 2609.13948首次发表:更新:

发表机构

Princeton University(普林斯顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过规则归纳任务评估推理模型的系统性,发现模型虽能解决原任务,但在结构等价变体上常失败,表明其行为缺乏系统性,难以稳健确立认知能力。

AI 中文摘要

人类认知的一个核心原则是系统性,即理解一个概念本质上与理解该概念的相近变体相关联。推理模型是否稳健地展现出这种系统性?如果是这样,我们应预期模型在同一任务的结构等价变体上表现一致。在此,我们扩展了认知科学中既定的规则归纳任务,以评估当前推理模型思维的系统性。每个任务族具有组合结构,我们利用该结构通过任务同构(如重组和替换)创建结构等价的任务变体。我们发现,尽管模型能够正确解决某个任务,但它们常常在同一任务的结构等价变体上失败。这些发现表明,许多模型行为缺乏系统性,这使得在特定评估情境之外稳健地确立推理模型的认知能力变得困难。

英文摘要

A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to understanding close variations of that concept. Do reasoning models robustly exhibit such systematicity? If so, we would expect consistent performance on structurally equivalent variants of the same task. Here, we extend established rule induction tasks from cognitive science to assess the systematicity of thought in current reasoning models. Each task family has compositional structure that we use to create structurally equivalent task variations through task isomorphisms such as recombination and substitution. We find that despite being able to correctly solve a task, models often fail on structurally equivalent variants of the same task. These findings suggest that many model behaviors lack systematicity, rendering it difficult to robustly establish the cognitive abilities of reasoning models beyond the particular contexts they were evaluated in.

CommentsCode available at https://github.com/smonsays/systematicity-eval

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑