arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ProAR:利用自回归视频模型学习前瞻性推理

ProAR: Learning Prospective Reasoning with Autoregressive Video Models

Linghui Shen, Tinghui Zhu, Sheng Zhang, Muhao Chen

arXiv 2610.03664首次发表:更新:

发表机构

The Hong Kong Polytechnic University; University of California, Davis; Microsoft(香港理工大学; 加利福尼亚大学戴维斯分校; 微软)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ProAR通过目标帧预测和未来表示自对齐,将自回归视频生成转为目标导向推理,以25%训练步骤超越标准基线,适用于具身推理。

AI 中文摘要

自回归(AR)视频模型在因果生成方面表现出色,但其对下一块预测的依赖使其局限于短视、反应式的范式。这一限制对以推理为导向的生成任务尤为关键,在这类任务中,通过有效的中间状态达成目标结果比局部视觉合理性更为重要。为应对这一挑战,我们提出了利用自回归视频模型学习前瞻性推理(ProAR),这是一种新颖的框架,将自回归视频生成转变为目标导向的推理过程。ProAR引入了两个关键组件:(1)为了将生成锚定到长期结果,我们通过非对称注意力掩码将目标帧预测集成到自回归循环中,使预测的目标帧能够引导中间状态的生成而不受其干扰。(2)为了引导短期转换,我们引入了未来表示自对齐机制,鼓励当前隐藏状态预判即将到来的时间动态。通过利用AR训练中的教师强制,我们在单次前向传播中提取干净的未来表示,并使用轻量级、仅训练时使用的预测器将当前表示与它们对齐。这两种机制共同将显式、稀疏的目标监督与隐式、密集的逐步指导无缝结合,以适度的计算成本促进连贯、目标导向的推理进展。实验表明,ProAR的互补组件在多种视觉推理基准上持续提升性能。该框架展现出极高的训练效率,仅使用25%的训练步骤即可超越完全训练的标准AR基线。这一范式还展示了在具身推理任务中的良好适用性。

英文摘要

Autoregressive (AR) video models excel at causal generation, but their reliance on next-chunk prediction confines them to a short-sighted, reactive paradigm. This limitation is particularly consequential for reasoning-oriented generation, where achieving a target outcome through valid intermediate states matters more than local visual plausibility. To address this challenge, we propose Learning Prospective Reasoning with Autoregressive Video Models (ProAR), a novel framework that transforms autoregressive video generation into a goal-oriented reasoning process. ProAR introduces two key components: (1) To anchor generation to the long-range outcome, we integrate goal-frame prediction into the autoregressive loop via an asymmetric attention mask, enabling the predicted goal frame to guide the generation of intermediate states without being disrupted by them. (2) To guide short-range transitions, we introduce future representation self-alignment to encourage current hidden states to anticipate upcoming temporal dynamics. By leveraging teacher-forcing in AR training, we extract clean future representations in a single forward pass and align current representations with them using a lightweight, training-only predictor. Together, these two mechanisms seamlessly combine explicit, sparse target supervision with implicit, dense step-wise guidance, promoting coherent, goal-directed reasoning progress with modest computational cost. Experiments show that ProAR's complementary components consistently improve performance across diverse visual reasoning benchmarks. The framework proves highly training-efficient, surpassing fully trained standard AR baselines using only 25% of the training steps. This paradigm also demonstrates promising applicability to embodied reasoning tasks.

CommentsProject Page: https://luka-group.github.io/ProAR/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑