arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向开放词汇三维场景理解的动态鲁棒光度-语义重建

Dynamic-Robust Photometric-Semantic Reconstruction for Open-Vocabulary 3D Scene Understanding

Boyu Cai, Li Yang, Yan Xu, Wei Liu, Nian Liu, Sikui Zhang, Yan Wang, Chunfeng Yuan, Weiming Hu

arXiv 2608.29177首次发表:更新:

发表机构

State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS); Institute of Automation, Chinese Academy of Sciences (CASIA); School of Information Science and Technology, ShanghaiTech University; The Chinese University of Hong Kong; Deepeleph Intelligent Technology(多模态人工智能系统国家重点实验室; 中国科学院自动化研究所; 上海科技大学信息科学与技术学院; 香港中文大学; 深智科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出SPAR架构与动态区域感知训练范式,解决NVS与OVS模型在动态场景的特征错位问题,在D-RE10K基准上实现SOTA性能,3/4视图下PSNR达22.15/23.33 dB,运动掩码预测mIoU达88.5%。

AI 中文摘要

近期,新颖视图合成(NVS)与开放词汇分割(OVS)的结合已产生强大的前馈三维基础模型。然而,这些模型固有地依赖静态场景假设,在不受约束的动态环境中会导致空间特征严重错位。为弥合这一关键差距,我们提出SPAR,一种新型联合语义-几何编码架构,其在潜在空间聚合前明确分离瞬态动态噪声。此外,我们引入动态区域感知的端到端训练范式,该范式将运动估计与多视图视觉及语义学习结构耦合。这种统一方法使网络能固有地解决运动冲突,并从动态输入中提取多视图一致、时间稳定的场景表示。在具有挑战性的D-RE10K基准上进行的大量实验表明,SPAR实现了最先进的性能。我们的端到端方法实现了卓越的新颖视图合成质量,仅使用3个和4个输入视图时,分别达到22.15 dB和23.33 dB的峰值信噪比(PSNR)。尽管以自监督方式训练,我们的模型在运动掩码预测上达到了88.5%的平均交并比(mIoU)。此外,我们的分析揭示了光度场景重建与语义理解之间存在强任务间协同作用,其中语义合成学习始终增强新颖视图渲染中的光度保真度。代码将在该httpsURL处提供。

英文摘要

The integration of novel view synthesis (NVS) and open-vocabulary segmentation (OVS) has recently yielded powerful feed-forward 3D foundation models. However, their inherent reliance on static-scene assumptions leads to severe misalignment of spatial features in unconstrained dynamic environments. To bridge this critical gap, we propose SPAR, a novel joint semantic-geometric encoding architecture that explicitly isolates transient dynamic noise prior to latent space aggregation. Furthermore, we introduce a dynamic-region-aware end-to-end training paradigm that structurally couples motion estimation with multi-view visual and semantic learning. This unified approach enables the network to inherently resolve motion conflicts and distill multi-view consistent, temporally stable scene representations from dynamic inputs. Extensive experiments on the challenging D-RE10K benchmark demonstrate that SPAR achieves state-of-the-art performance. Our end-to-end approach achieves exceptional novel view synthesis quality, yielding a PSNR of 22.15 dB and 23.33 dB given only 3 and 4 input views respectively. Despite being trained in a self-supervised manner, our model achieves an mIoU of 88.5% for motion mask prediction. Furthermore, our analysis reveals a strong inter-task synergy between photometric scene reconstruction and semantic understanding, where semantic synthesis learning consistently enhances photometric fidelity in novel view rendering. Code will be available at https://github.com/dmucby/SPAR.

CommentsAccepted to the European Conference on Computer Vision (ECCV) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑