arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VBVR-Pro:用于原生视觉推理的可扩展且可验证工具套件

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Junxiang Xu, Ruisi Wang, Fanyi Pu, Maijunxian Wang, Ran Ji, Tongxi Zhou, Chenyang Gu, Jing Zuo, Hongcan Xiao, Yimeng Geng, Wanqi Yin, Wei Chen, Oscar Qian, Zhengan Yan, Ziqi Huang, Haiwen Diao, Liang Pan, Bo Li, Xiangyu Fan, Dezhi Luo, Fengyuan Yu, Zehong Zhao, Qingying Gao, Tinghui Zhu, Yilan Zhang, Jingqi Tong, Pinyuan Feng, Zhengze Jiang, Letian Wang, Ziyu Guo, Renrui Zhang, Jieneng Chen, Sonia Joseph, Constantin Venhoff, Saman Motamed, Mengyue Yang, Chandra Sripada, Alan Yuille, Philip Torr, Lvmin Zhang, Vikash Kumar, Daniel Khashabi, Nikolaus Kriegeskorte, Raphaël Millière, Vincent C. Müller, Anyi Rao, Quan Wang, Ziwei Liu, Dahua Lin, Lei Yang, Hokin Deng, Zhongang Cai

arXiv 2608.26105首次发表:更新:

发表机构

Nanyang Technological University; University of California, Berkeley; University of California, San Diego; The University of Tokyo; The Chinese University of Hong Kong; University of Michigan; Johns Hopkins University; University of California, Davis; University of California, Los Angeles; Carnegie Mellon University; Columbia University; University of Toronto(南洋理工大学; 加州大学伯克利分校; 加州大学圣地亚哥分校; 东京大学; 香港中文大学; 密歇根大学; 约翰霍普金斯大学; 加州大学戴维斯分校; 加州大学洛杉矶分校; 卡内基梅隆大学; 哥伦比亚大学; 多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出VBVR-Pro,一个可扩展可验证的原生视觉推理闭环测试平台,含300个生成任务、确定性评分器及多生成器机制研究,训练模型在多视觉推理基准上迁移性强,发布全部相关资源。

AI 中文摘要

原生视觉推理将视觉生成本身视为推理的媒介:视觉状态(即图像和视频)不仅是需要理解的输入或需要渲染的输出,更是超越语言进行问题解决的一级载体。然而,进展仍受到可扩展训练任务、可靠反馈以及跨生成载体的受控比较的缺乏所阻碍。本研究中,我们引入VBVR-Pro,这是一个闭环测试平台,它使通过生成实现的原生视觉推理具备可训练、可验证、可优化和实验可控的特性。1)任务扩展:VBVR-Pro将视觉推理转化为包含300个程序生成任务的受控任务空间。在VBVR-Pro上训练的模型在超出该套件的7个外部视觉推理基准(如RISE-Video、MME-CoF-Pro和BabyVision)上表现出强迁移能力。2)可验证奖励:VBVR-Pro为基于任务的评估提供可验证的奖励评分器。通过对领先多模态大语言模型(MLLM)作为评判的系统研究,我们确定了流行的视觉语言模型(VLM)作为评判范式的反复出现的失败模式。相比之下,所提出的评分器基于确定性、任务特定规则,实现了与人类判断的细粒度对齐。重要的是,它们可作为大规模多任务强化学习的可靠奖励信号,并在视觉推理任务中表现出更强的强化学习后性能。3)机制研究:VBVR-Pro支持对超过30种图像、视频及交错生成器进行受控模态研究。我们的分析表明,对于需要持续时空状态跟踪的任务,视频生成仍然表现最强,而交错生成提供了计算高效的替代方案。关键的是,消融研究和探测实验表明,存在对视觉推理至关重要的原生视觉轨迹。我们发布所有数据、模型、评分器和代码。

英文摘要

Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.

CommentsHomepage: https://video-reason.com/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑