arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Dyn-3D:揭示与解决视觉语言模型中的自身运动歧义

Dyn-3D: Unveiling and Resolving Ego-Motion Ambiguity in Vision-Language Models

Jiayu Ding, Zhuodong Liu, Lei Zhang, Manyu Xiong, Hongbo Jin, Haoran Tang, Hongbo Zhang, Changen Zhu, Wenbo Xing

arXiv 2609.01059首次发表:更新:

发表机构

InkMind Team(InkMind团队)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉语言模型处理动态3D空间推理时出现的运动学崩溃问题,引入Dyn-3D基准,提出含Kinematic-GSPO算法的TempoVista框架,提升了运动估计与鲁棒空间推理能力。

AI 中文摘要

当视觉语言模型(VLMs)处理动态3D空间推理时,自身运动感知对于解决单目尺度歧义至关重要。然而,当前模型常常过拟合于平滑轨迹先验,而非真正理解物理运动,导致在大位移下空间推理严重退化,我们将这一现象称为运动学崩溃。该问题源于自然视频中虚假的视觉-运动相关性以及缺乏显式物理监督。为评估此问题,我们引入Dyn-3D基准,采用反事实3D渲染严格解耦视觉变化与真实运动学属性。此外,我们提出TempoVista框架,包含Kinematic-GSPO算法,通过将度量物理真值嵌入策略优化,使TempoVista将视觉表示明确锚定在3D空间中。实验表明,我们的方法利用相机动力学作为有效几何校准信号,显著提升了运动估计与鲁棒空间推理能力。

英文摘要

As Vision-Language Models (VLMs) tackle dynamic 3D spatial reasoning, ego-motion perception becomes essential to resolve monocular scale ambiguity. However, current models often overfit to smooth trajectory priors rather than genuinely understanding physical motion. Consequently, their spatial reasoning degrades severely under large displacements, a phenomenon we term Kinematic Collapse. This failure stems from spurious visual-motion correlations in natural videos and a lack of explicit physical supervision. To evaluate this, we introduce Dyn-3D, a benchmark using counterfactual 3D rendering to rigorously decouple visual changes from true kinematic properties. Furthermore, we propose the TempoVista framework, featuring the Kinematic-GSPO algorithm. By embedding metric physical ground truth into policy optimization, TempoVista explicitly grounds visual representations in 3D space. Experiments demonstrate that our approach significantly improves both motion estimation and robust spatial reasoning by utilizing camera dynamics as an effective geometric calibration signal.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑