arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在类ARC任务中利用测试时任务嵌入进行隐式规则归纳

Implicit Rule Induction with Test-Time Task Embeddings in ARC-like Tasks

Adrien Deliège, Claas Beger, Marc Van Droogenbroeck, Melanie Mitchell

arXiv 2609.21181首次发表:更新:

发表机构

University of Liège; Santa Fe Institute(列日大学; 圣塔菲研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出两步测试时训练协议Embed-TTT,通过先微调任务嵌入再微调骨干网络,在ARC类任务中实现更好的隐式规则归纳,提升嵌入质量与检索性能,并支持规则插值。

AI 中文摘要

抽象与推理语料库及相关基准测试评估AI模型能否解决新颖的推理任务,但往往不清楚成功是反映了对预期底层规则的推断,还是依赖于捷径。我们通过研究Vision ARC(VARC)中的测试时任务嵌入来解决这一空白,VARC是一种模型,其中预训练骨干网络辅以一个可训练的任务嵌入,该嵌入表示变换规则。在原始VARC中,测试时训练(TTT)被联合应用于骨干网络和任务嵌入。在此,我们引入一种新颖的两步TTT协议:首先仅微调任务嵌入(Embed-TTT),然后冻结它并微调骨干网络。在ARC-AGI-1、ConceptARC以及两个具有已知规则的控制数据集上,Embed-TTT一致地产生改进的任务嵌入,这些嵌入与底层任务规则更好地对齐,改善了基于嵌入的检索,并能够对已知规则进行准确的线性探测。定性上,Embed-TTT在ARC-AGI-1上识别出测试任务与训练任务之间更多语义上有意义的关系。我们还表明,仅优化任务嵌入(不到模型参数的0.01%)就已经解决了ARC-AGI-1、ConceptARC和Mini-ARC任务中不可忽视的一部分,而完整的两步流程则提高了最终性能。最后,我们表明Embed-TTT恢复了参数化规则的底层几何结构,并学习了能够实现规则级插值(而非外推)的组合能力。这些发现支持在类ARC评估中更清晰地区分规则归纳与规则执行,激励了更好地区分分布内与分布外规则的基准测试。

英文摘要

The Abstraction and Reasoning Corpus and related benchmarks evaluate whether AI models can solve novel reasoning tasks, but often leave unclear whether success reflects inference of the intended underlying rule or reliance on shortcuts. We address this gap by studying test-time task embeddings in Vision ARC (VARC), a model in which a pre-trained backbone is complemented by a trainable embedding representing the transformation rule. In the original VARC, test-time training (TTT) is jointly applied to the backbone and task embedding. Here we introduce a novel two-step TTT protocol: first finetune only the task embedding (Embed-TTT), then freeze it and finetune the backbone. Across ARC-AGI-1, ConceptARC, and two controlled datasets with known rules, Embed-TTT consistently yields improved task embeddings, ones that align better with underlying task rules, improve embedding-based retrieval, and enable accurate linear probing of known rules. Qualitatively, Embed-TTT identifies more semantically meaningful relations between test and train tasks on ARC-AGI-1. We also show that optimizing only task embeddings (less than 0.01% of model parameters) already solves a non-trivial fraction of ARC-AGI-1, ConceptARC, and Mini-ARC tasks, while the full two-step pipeline improves final performance. Finally, we show that Embed-TTT recovers the underlying geometric structure of parametric rules and learns compositional capabilities that enable rule-wise interpolation, but not extrapolation. These findings support a clearer separation between rule induction and rule execution in ARC-like evaluations, motivating benchmarks that better distinguish in-distribution from out-of-distribution rules.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑