展开世界:分解四维属性以增强空间推理能力
Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning
- Joy Future Academy(京东探索研究院)
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- The Hong Kong University of Science and Technology(香港科技大学)
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对VLMs的空间推理瓶颈,提出分解式强化学习框架FactoSR,将世界一致性推理分解为三个几何子目标,在相关基准上实现三维、四维推理性能的显著提升,为发展具世界感知的VLMs提供关键路径。
AI中文摘要:
尽管视觉-语言模型(VLMs)在通用多模态任务中表现出卓越的能力,但它们在对物理世界进行推理时仍本质上是“扁平的”。我们认为,这种空间瓶颈源于深刻的维度不匹配:VLMs被训练用于解释二维投影,而真正的空间推理需要恢复潜在的三维几何结构和时间连续性。为了克服这种高维复杂性,我们倡导从整体学习转向“分而治之”的范式。我们提出FactoSR,这是一种分解式强化学习框架,可明确解释视觉投影所坍缩的维度。FactoSR的核心是将世界一致性推理的整体问题分解为三个正交的几何子目标:平面对应关系(XY)、深度一致性(Z)和时间可逆性(T)。通过在统一的策略学习机制内优化这些可验证的约束,我们有效地将不适定的投影恢复问题转化为一系列切实的推理步骤。在多视图和视频基准上的广泛评估表明,这种优雅的分解在三维和四维推理中取得了显著提升,在VSI-Bench上实现了5.9%的提升,在All-Angles-Bench上实现了4.5%的提升。我们的发现表明,增强显式的分解式四维一致性是将VLMs发展为鲁棒的、具有世界感知能力的推理器的关键一步。
英文摘要:
Despite the remarkable prowess of Vision-Language Models (VLMs) in general multimodal tasks, they remain fundamentally ``flat'' when reasoning about the physical world. We argue that this spatial bottleneck stems from a profound dimensional mismatch: while VLMs are trained to interpret 2D projections, true spatial reasoning demands the recovery of latent 3D geometry and temporal continuity. To conquer this high-dimensional complexity, we advocate a shift from monolithic learning to a ``divide and conquer'' paradigm. We present FactoSR, a factorized reinforcement learning framework that explicitly interpret the dimensions collapsed by visual projection. At its core, FactoSR decomposes the monolithic problem of world-consistent reasoning into three orthogonal, geometric sub-objectives: planar correspondence ($XY$), depth consistency ($Z$), and temporal reversibility ($T$). By optimizing these verifiable constraints within a unified policy learning mechanism, we effectively transform an ill-posed projection recovery problem into a series of tangible reasoning steps. Extensive evaluations on multi-view and video benchmarks demonstrate that this elegant decomposition yields substantial gains in 3D and 4D reasoning, achieving a 5.9% boost on VSI-Bench and 4.5% on All-Angles-Bench. Our findings suggest that reinforcing explicit, factorized 4D consistency is a critical step toward evolving VLMs into robust, world-aware reasoners.