AI 中文总结
本研究证明冻结视觉模型的潜在向量算术可隐式推理规则,无需微调,在Bongard和ARC任务上接近或超越基线,直觉向量实现高准确率检索与任务识别。
AI 中文摘要
大型自监督视觉模型学习到的表示能够支持场景分割和物理对象的语义分解。我们探究其表示几何是否能在无需任何任务特定微调的情况下,迁移到视觉推理问题。我们假设关系表示可能通过编码视觉输入之间的相似性和变换,从而桥接感知和抽象推理,使得一种直觉的、隐式的推理形式可以通过潜在向量算术执行。具体而言,我们在抽象和自然主义的Bongard问题、ARC-AGI-1和ARC-AGI-2以及新提出的ARC-GEN实例上检验了DINOv3、MAE和随机像素投影。在两个Bongard基准上,冻结视觉嵌入的简单最近质心读取的准确率与原始基准报告的任务特定基线相差在四个百分点以内。在ARC中,总结演示输入-输出变换的潜在差异向量(我们称之为直觉向量)与来自同一任务的测试向量显示出更高的一致性,而来自无关任务的向量则近乎正交。这种潜在几何是可操作的:沿其直觉向量传输查询持续提高了精确输出检索,在ARC-AGI-2评估上达到70.7。在来自794个任务的397,000个ARC-GEN实例中,单对直觉向量以约87%的留一法准确率识别生成任务。这些发现表明,对冻结视觉表示进行潜在向量算术支持跨不同问题领域的隐式规则推断,而无需生成模型组件,这表明推断抽象变换和生成其特定实例的结果可能是可分离的能力。
英文摘要
Large self-supervised vision models learn representations that support scene segmentation and the semantic decomposition of physical objects. We ask whether their representational geometry supports transfer to visual reasoning problems without any task-specific fine-tuning. We hypothesized that relational representations may bridge perception and abstract reasoning by encoding similarities and transformations among visual inputs such that an intuitive, implicit form of reasoning may be performed via latent vector arithmetic. Specifically, we examine DINOv3, MAE, and random pixel projections on abstract and naturalistic Bongard problems, ARC-AGI-1 and ARC-AGI-2, and novel ARC-GEN instances. On both Bongard benchmarks, the accuracy of a simple nearest-centroid readout of frozen visual embeddings is within four percentage points of the task-specific baselines reported with the original benchmarks. In ARC, latent difference vectors summarizing demonstration input-output transformations, which we refer to as intuition vectors, show greater alignment with test vectors from the same task, whereas those from unrelated tasks are near orthogonal. This latent geometry is operational: transporting a query along its intuition vector consistently improves exact-output retrieval, reaching 70.7 on ARC-AGI-2 evaluation. Across 397,000 ARC-GEN instances from 794 tasks, single-pair intuition vectors identify the generating task with approximately 87\% leave-one-out accuracy. These findings suggest that latent vector arithmetic over frozen visual representations supports implicit rule inference across varied problem domains without a generative model component, indicating that inferring an abstract transformation and generating its instance-specific consequence may be separable capacities.