arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当前系统性泛化任务遗漏了什么?一种以推理为中心的分析

What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered Analysis

Chengwen Qi, Deheng Ye, Yatao Bian

arXiv 2609.19212首次发表:更新:

发表机构

National University of Singapore; Nanyang Technological University(新加坡国立大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对现有系统性泛化任务过度简化的问题,提出融合演绎、归纳与溯因推理的TranSGrid测试平台,实验表明现有任务降低了归纳或溯因需求,全面评估需同时涵盖三种推理。

AI 中文摘要

系统性泛化,即通过重新组合已知的原子元素来解决新问题的能力,是人类智能的核心,但在受控条件下难以严格研究。因此,现有研究依赖于诸如近似线性动作组合、基于生产率的测试和动作显式目标等简化手段,这些简化使系统性泛化更易于研究,但遗漏了这一能力的某些基本方面。为了刻画这些简化遗漏了什么,我们采用以推理为中心的视角,并引入TranSGrid,一个在统一任务中结合演绎、归纳和溯因推理的测试平台。在4,800个TranSGrid实例上对七个Transformer进行的实验表明,所有模型在TranSGrid上的表现远差于在保留测试集上的表现:最大的模型解决了测试集的79.6%,但仅解决了TranSGrid的55.3%和最困难子集的15.8%。这一差距在训练长度范围内仍然存在,表明仅靠生产率不足以评估系统性泛化。此外,我们将另外两种简化重新引入TranSGrid:一种变体使动作几乎线性组合(降低归纳需求),另一种使目标动作显式(降低溯因需求)。在两种情况下,解决率都大致恢复到测试集水平,表明任一简化单独就足以将TranSGrid降低为普通的保留测试集。综合来看,我们的结果表明,现有任务降低了归纳和溯因需求中的一种或两种,而全面衡量系统性泛化需要一个涉及所有三种推理形式的任务。

英文摘要

Systematic generalization, the ability to solve novel problems by recombining known atomic elements, is central to human intelligence but difficult to study rigorously under controlled settings. Existing studies therefore rely on simplifications such as elemental composition, productivity-based tests, and action-explicit goals, which make systematic generalization easier to study but omit some essential aspects of this capability. To characterize what these simplifications miss, we adopt a reasoning-centered lens and introduce TranSGrid, a testbed that brings deductive, inductive, and abductive reasoning together within a unified task. Experiments with seven Transformer models on 4,800 TranSGrid instances show that all models perform much worse on TranSGrid than on a held-out test set: the largest model solves 79.6% of the test set, but only 55.3% of TranSGrid and 15.8% of the hardest subset. The gap remains within the training length range, showing that productivity alone is not sufficient to evaluate systematic generalization. Additionally, we reintroduce the other two simplifications into TranSGrid: one variant limits interactions among action effects to approximate elemental composition (reducing the inductive demand); the other makes goals action-explicit (reducing the abductive one). In both, solve rates return to roughly the test set level, showing that either simplification alone is enough to reduce TranSGrid to an ordinary held-out test set. Together, our results show that existing tasks reduce either or both of the inductive and abductive demands of systematic generalization, and that comprehensively measuring this capability requires a task that involves all three forms of reasoning.

CommentsPreprint. Under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑