arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12800cs.CV

UniVR:用于统一视觉推理的视觉空间思考

UniVR: Thinking in Visual Space for Unified Visual Reasoning

Zhongwei Ren, Yunchao Wei, Yao Zhao, Weibo Gong, Xiao Liu, Anran Wang, Xiangtai Li, Xiaojie Jin

首次发表
浏览论文内容

中文总结 AI 辅助

UniVR旨在从纯视觉演示中同时学习复杂推理、物理动力学和长期规划,核心是VR-GRPO强化学习范式。构建VR-X基准进行训练和评估,在VR-X上有25%的提升,还提高了多模态理解基准性能,相关资源已开源。

中文摘要 AI 辅助

从原始视觉数据中直接学习广泛的世界知识是智能的一项基本能力。我们引入了UniVR,首次探索从纯视觉演示中同时学习复杂推理、细粒度物理动力学和长期规划。其核心是VR-GRPO,一种具有互补全局和步级奖励的强化学习范式。该方法在整个推理过程中强制逻辑连贯和物理一致性,无需特定任务启发式或图像-文本对。为训练和评估UniVR,我们构建了VR-X,这是一个从16个不同来源策划的大规模基准,涵盖长期操纵、空间谜题和物理推理。它是首个在纯视觉协议下评估这些异构能力的综合套件。值得注意的是,UniVR在VR-X上实现了高达25%的提升,其卓越的视觉推理还提高了各种多模态理解基准的性能。这些发现强调了视觉空间内推理的巨大潜力,所有代码、数据和模型都已开源以供进一步研究。

英文摘要

Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations. At its core, UniVR features VR-GRPO, a reinforcement learning paradigm with complementary global and step-level rewards. This approach enforces logical coherence and physical consistency throughout the reasoning process without requiring task-specific heuristics or image-text pairs. To train and evaluate UniVR, we construct VR-X, a large-scale benchmark curated from 16 diverse sources spanning long-horizon manipulation, spatial puzzles, and physical reasoning. It is the first comprehensive suite to assess these heterogeneous capabilities under a purely visual protocol. Remarkably, UniVR achieves up to a 25% improvement on VR-X, and its superior visual reasoning also boosts performance on various multimodal understanding benchmarks. These findings underscore the vast potential of reasoning within visual spaces, with all code, data, and models are open-sourced for further research.

发表机构

  • Beijing Jiaotong University(北京交通大学)
  • ByteDance(字节跳动)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑