arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Scene-SAM3D:无需微调的多视图场景资产生成

Scene-SAM3D: Multi-View Scene Asset Generation Without Fine-Tuning

Yuqi Zhang, Yadan Luo, Xiangyu Sun, Fengyi Zhang, Zi Huang, Xin Tan

arXiv 2607.16805首次发表:更新:

发表机构

The University of Queensland; East China Normal University; Shanghai AI Laboratory(昆士兰大学; 华东师范大学; 上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现实场景中3D生成难题,提出无需训练的Scene-SAM3D框架,通过选互补视图、潜在速度融合及高斯优化,实现多视图场景资产生成,实验证明其在减少场景级CD和降低计算成本上有显著成效。

AI 中文摘要

高质量3D场景资产对机器人操作、导航和模拟等应用至关重要。尽管单图像3D生成模型有强大的对象先验,但在现实场景中仍不足。本文介绍Scene-SAM3D,一个无需训练的框架,将SAM3D从单视图对象生成扩展到校准多视图场景资产生成。它选择紧凑的互补视图集,减少观察冗余并为遮挡区域提供额外证据。基于所选视图,进行高效的潜在速度融合以整合多视图证据并抑制规范空间中的跨视图冲突。最后,通过轻量级刚性对象高斯优化在200次迭代内优化场景布局。在Replica和ScanNet++上的实验表明,该方法在实例和场景级别都有持续改进,在相同多视图设置下,减少了场景级CD,还降低了流模型采样FLOP和运行时间延迟。代码将在指定网址发布。

英文摘要

High-quality 3D scene assets are critical for embodied applications such as robotic manipulation, navigation, and simulation. Despite their strong object priors, recent single-image 3D generation models such as SAM3D remain insufficient for real-world scenes, where severe occlusions, redundant observations, and cross-view inconsistencies make reliable scene generation challenging. We introduce Scene-SAM3D, a training-free framework that extends SAM3D from single-view object generation to calibrated multi-view scene asset generation. Scene-SAM3D selects a compact set of complementary views, reducing observation redundancy while providing additional evidence for regions occluded in individual views. Based on the selected views, it performs step-efficient latent velocity fusion to integrate multi-view evidence and suppress cross-view conflicts in canonical space. Finally, a lightweight rigid-object Gaussian optimization refines the scene layout within 200 iterations while preserving the generated object geometry. Experiments on Replica and ScanNet++ demonstrate consistent improvements at both instance and scene levels, with our method reducing scene-level CD by 43.8% on Replica and 30.9% on ScanNet++, while cutting flow-model sampling FLOPs and wall-time latency by nearly 20% under the same multi-view setting. Code will be released at https://github.com/xibi777/Scene-SAM3D.

Comments18 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑