发表机构
Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SAMV-DUSt3R,将SAM2掩码注入MV-DUSt3R,通过交叉流掩码块和空间RankGNN实现实例级三维解耦,精度提升11%。
AI 中文摘要
随着从三维场景中解耦物体的需求日益增长,我们提出了SAMV-DUSt3R,一种端到端模型,将SAM2二维掩码注入到MV-DUSt3R重建中。一个交叉流掩码块利用这些掩码引导网络朝向目标实例,共同提高形状精度并实现物体级解耦,无需多阶段流水线。为确保重建稳定性,一个轻量级的空间RankGNN以73.5%的选择准确率选择最优参考视图。大量实验表明,与最先进的基线相比,我们的方法在各种指标上将平均重建精度提高了11%。这些结果揭示了强大的实例解耦能力,并对驾驶、机器人、AR/VR和遗产数字化有明确益处。
英文摘要
With the rising demand to decouple objects from 3D scenes, we propose SAMV-DUSt3R, an end-to-end model that injects SAM2 2D masks into MV-DUSt3R reconstruction. A Cross Flow Mask Block uses these masks to steer the network toward the target instance, jointly improving shape accuracy and achieving object-level disentanglement without multi-stage pipelines. To ensure reconstruction stability, a lightweight Spatial RankGNN selects the optimal reference view with a selection accuracy of 73.5\%. Extensive experiments demonstrate that our method boosts average reconstruction precision by 11\% across various metrics compared to state-of-the-art baselines. These results reveal a strong instance-disentanglement capability and clear benefits for driving, robotics, AR/VR, and heritage digitisation.