arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TraceViT:用于视觉抽象推理的基于轨迹的监督

TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning

Binnan Liu, Yechi Ma, Tian Xie, Wei Hua

arXiv 2607.29586首次发表:更新:

AI 中文总结

针对视觉抽象推理任务,提出基于语义单调变换链监督的循环视觉推理器TraceViT,在ARC-AGI-1、ARC-AGI-2上取得对应准确率,证实轨迹监督需与约束配对才有效。

AI 中文摘要

抽象与推理语料库(ARC)测试模型能否从少量输入-输出示例中推断出未见过的变换并将其应用于新的网格。循环视觉推理器会在多次迭代中细化预测,但传统训练仅约束最终输出,未约束中间细化过程。我们认为这些细化过程应逐步遵循变换。我们引入TraceViT,一种用语义单调变换链训练的循环视觉推理器,这些链通过重写和验证程序化任务实现得到,将每个解决方案分解为中间网格状态。每次迭代由少量演示得到的任务参考和表示当前网格状态的对象工作空间进行约束。由于这些链的长度可能与循环不同,软轨迹对齐仅约束它们的顺序,让模型自由分配迭代次数。TraceViT在ARC-AGI-1上达到67.8%的pass@2,在ARC-AGI-2上达到24.3%。对ARC-AGI-1的受控消融实验显示,轨迹监督仅在与约束配对时才有益。代码和数据将在此https URL上提供。

英文摘要

The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-output examples and apply it to a new grid. Looped visual reasoners refine predictions over multiple iterations, but conventional training constrains only the final output, leaving intermediate refinements unconstrained. We propose that these refinements should instead follow the transformation step by step. We introduce TraceViT, a looped visual reasoner trained with semantically monotonic transformation chains. We obtain these chains by rewriting and verifying programmatic task implementations, decomposing each solution into intermediate grid states. Each iteration is grounded by a task reference derived from the few-shot demonstrations and an object workspace representing the current grid state. Because these chains may differ in length from the loop, soft trace alignment enforces only their ordering, letting the model allocate iterations freely. TraceViT achieves 67.8% pass@2 on ARC-AGI-1 and 24.3% on ARC-AGI-2. Controlled ablations on ARC-AGI-1 show that trace supervision becomes beneficial only when paired with grounding. Code and data will be available at https://github.com/LiuBinnan/TraceViT.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑