arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01286cs.CV

Dyna3:基于深度基础模型的VLM引导免训练4D重建

Dyna3: VLM-Guided Training-Free 4D Reconstruction via Depth Foundation Models

Xinhao Xiang, Weiyang Li, Zhijie Zheng, Abhijeet Rastogi, Jiawei Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

Dyna3利用深度基础模型DA3和VLM引导的SAM 3分割,无需训练即可实现4D动态场景重建,在分割精度、速度和内存上均优于现有方法。

中文摘要 AI 辅助

最近的深度基础模型如Depth Anything 3(DA3)在多视角深度估计上取得了显著成果,但假设场景为静态3D,限制了其在真实动态环境中的适用性。现有的免训练4D方法如Easi3R和VGGT4D依赖于经过对应关系训练的骨干网络,其注意力编码跨帧匹配,而这一特性在仅深度模型中如DA3中缺失。我们提出Dyna3,一个无需微调的免训练框架,扩展DA3以进行4D动态场景重建。我们的关键洞察是,DA3的跨视角特征虽然仅针对深度一致性训练,但结合跨帧最佳匹配特征搜索时,隐式编码了运动判别信号。其静态表面能全局找到一致匹配,而动态对象则不能。我们进一步采用视觉语言模型(VLM)自动生成场景特定的语义提示用于SAM 3,实现精确的实例级分割,区分哪些对象移动以及存在哪些对象。对于重建,我们将场景解耦为跨帧对齐的静态背景和逐帧动态点云。在四个数据集上的实验表明,Dyna3在动态对象分割上超过基于对应关系训练的方法,比最先进的VGGT4D高出+5.5pp J-Mean,同时实现高达13倍更快的姿态估计和3倍更快的4D重建,内存降低4到8倍。因此,Dyna3能够支持先前方法无法支持的更密集的时间采样。

英文摘要

Recent depth foundation models like Depth Anything 3 (DA3) achieve remarkable multi-view depth estimation but assume static 3D scenes, limiting their applicability to real-world dynamic environments. Existing training-free 4D methods like Easi3R and VGGT4D rely on correspondence-trained backbones whose attention encodes cross-frame matching, a property absent in depth-only models like DA3. We present Dyna3, a training-free framework that extends DA3 for 4D dynamic scene reconstruction without any fine-tuning. Our key insight is that DA3's cross-view features, though trained only for depth consistency, implicitly encode motion-discriminative signals when combined with best-match feature search across frames. Its static surfaces find consistent matches globally, while dynamic objects cannot. We further adopt vision-language models (VLM) to automatically generate scene-specific semantic prompts for SAM 3, enabling precise instance-level segmentation that distinguishes which objects move from what objects exist. For reconstruction, we decouple the scene into a cross-frame aligned static background and per-frame dynamic point clouds. Experiments on four datasets demonstrate that Dyna3 surpasses correspondence-trained methods with +5.5pp J-Mean over state-of-the-art VGGT4D on dynamic object segmentation, while achieving up to 13x faster pose estimation and 3x faster 4D reconstruction with 4 to 8x lower memory. Dyna3 could therefore enable much denser temporal sampling that prior methods cannot support.

发表机构

  • University of California, Davis(加州大学戴维斯分校)
  • Genies Inc.(Genies公司)

机构由 AI 辅助整理,请以论文原文为准。

↑