发表机构
Shenzhen Campus of Sun Yat-sen University; Shenzhen Loop Area Institute; Tsinghua Shenzhen International Graduate School; Tencent; Mohamed bin Zayed University of Artificial Intelligence; UiT The Arctic University of Norway(中山大学深圳校区; 深圳河套学院; 清华大学深圳国际研究生院; 腾讯; 穆罕默德·本·扎耶德人工智能大学; 挪威北极大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
VR-JEPA通过局部对比状态学习对齐V-JEPA预测器与任务推理逻辑,利用预测的潜在轨迹引导视频生成,在VBVR-Pro-Bench上相对提升11.33%,减少物理伪影并增强逻辑一致性。
AI 中文摘要
通过视频生成进行推理,通过建模潜在视觉状态及其动态,为视觉智能提供了一条有前景的路径。然而,当前的视频生成模型通常缺乏关于这些状态应如何演变的明确指导,导致生成的轨迹容易出现物理和结构上的不一致,从而削弱推理的可靠性。虽然视频联合嵌入预测架构(V-JEPA)通过潜在预测提供了丰富的时空先验,但这些通用先验并不能自然地适应复杂视觉任务所需的逻辑推理能力。为弥合这一差距,我们提出了VR-JEPA,一种通过局部对比状态学习将V-JEPA预测器与任务特定推理逻辑对齐的框架,并利用其预测的潜在轨迹来指导视频生成以实现视觉推理。具体而言,(i)我们在相同输入条件下将成功轨迹与生成的替代轨迹配对,并利用它们在V-JEPA表示中的差异来识别信息丰富的状态和标记,以进行局部对比监督。(ii)我们进一步为V-JEPA预测器配备在锚定任务数据上训练的特定技能专家,使模型能够在其共享的时空先验基础上,跨不同认知领域自适应地专业化。结合特定技能专家,这种对比监督使VR-JEPA能够预测提供任务特定逻辑指导的潜在轨迹,以用于视频生成。在大型VBVR-Pro-Bench数据集上的全面实验表明,VR-JEPA相对于最先进的基于生成的推理基线实现了11.33%的相对改进,显著减少了物理伪影并增强了逻辑一致性。
英文摘要
Reasoning through video generation offers a promising path toward visual intelligence by modeling latent visual states and their dynamics. However, current video generation models often lack explicit guidance on how these states should evolve, leaving generated trajectories prone to physical and structural inconsistencies that undermine reasoning reliability. While the Video Joint-Embedding Predictive Architecture (V-JEPA) provides rich spatiotemporal priors learned through latent prediction, these general priors do not naturally adapt to the logical reasoning capabilities required for complex visual tasks. To bridge this gap, we propose VR-JEPA, a framework that aligns the V-JEPA predictor with task-specific reasoning logic through localized contrastive-state learning and uses its predicted latent trajectories to guide video generation for visual reasoning. Specifically, (i) we pair successful trajectories with generated alternatives under the same input conditions and use discrepancies in their V-JEPA representations to identify informative states and tokens for localized contrastive supervision. (ii) We further equip the V-JEPA predictor with skill-specific experts trained on anchor-task data, allowing the model to adaptively specialize its shared spatiotemporal priors across diverse cognitive domains. Together with skill-specific experts, this contrastive supervision enables VR-JEPA to predict latent trajectories that provide task-specific logical guidance for video generation. Comprehensive experiments on the large-scale VBVR-Pro-Bench dataset demonstrate that VR-JEPA achieves an $11.33\%$ relative improvement over the cutting-edge generation-based reasoning baseline, significantly mitigating physical artifacts and enhancing logical consistency.