arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAMV-DUSt3R:从稀疏多视图进行实例中心的三维场景解耦

SAMV-DUSt3R: Instance-Centric 3D Scene Decoupling from Sparse Multi-Views

Langxu Zhao, Zuan Gu, Yingdan Zhang, Pengfei Zhao, Tianhan Gao

arXiv 2609.11279首次发表:更新:

发表机构

Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SAMV-DUSt3R,将SAM2掩码注入MV-DUSt3R,通过交叉流掩码块和空间RankGNN实现实例级三维解耦,精度提升11%。

AI 中文摘要

随着从三维场景中解耦物体的需求日益增长,我们提出了SAMV-DUSt3R,一种端到端模型,将SAM2二维掩码注入到MV-DUSt3R重建中。一个交叉流掩码块利用这些掩码引导网络朝向目标实例,共同提高形状精度并实现物体级解耦,无需多阶段流水线。为确保重建稳定性,一个轻量级的空间RankGNN以73.5%的选择准确率选择最优参考视图。大量实验表明,与最先进的基线相比,我们的方法在各种指标上将平均重建精度提高了11%。这些结果揭示了强大的实例解耦能力,并对驾驶、机器人、AR/VR和遗产数字化有明确益处。

英文摘要

With the rising demand to decouple objects from 3D scenes, we propose SAMV-DUSt3R, an end-to-end model that injects SAM2 2D masks into MV-DUSt3R reconstruction. A Cross Flow Mask Block uses these masks to steer the network toward the target instance, jointly improving shape accuracy and achieving object-level disentanglement without multi-stage pipelines. To ensure reconstruction stability, a lightweight Spatial RankGNN selects the optimal reference view with a selection accuracy of 73.5\%. Extensive experiments demonstrate that our method boosts average reconstruction precision by 11\% across various metrics compared to state-of-the-art baselines. These results reveal a strong instance-disentanglement capability and clear benefits for driving, robotics, AR/VR, and heritage digitisation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑