arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STAR:面向通用3D场景理解的空间拓扑感知路由框架

STAR: A Spatial-Topology Aware Routing Framework for Generalizable 3D Scene Understanding

Mingwei Xing, Xinliang Wang, Yifeng Shi

arXiv 2608.11699首次发表:更新:

发表机构

KE Holdings Inc.(贝壳找房)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对传感器模态拓扑差异导致的统一3D场景理解模型构建难题,提出STAR框架,引入多属性自监督预训练分支与DSR、EDA机制,在ScanNet、S3DIS等数据集上取得优于基线的性能。

AI 中文摘要

构建统一的3D场景理解模型长期受到传感器模态间拓扑差异的阻碍。尽管应用混合专家(Mixture-of-Experts,MoE)架构是一种灵活的多领域3D理解方法,但我们观察到,仅基于特征的传统MoE路由在语义监督下可能无法充分表征局部采样拓扑,当语义一致性与几何异质性共存时,会导致专家分配困难。为克服这一挑战,我们提出STAR(Spatial-Topology Aware Routing Framework,空间拓扑感知路由框架)。具体而言,我们引入了一个多属性自监督预训练分支,涵盖拓扑和纹理变化,以锚定跨领域结构先验。在此基础上,我们设计了一个领域感知专家分支,包含两种机制:领域空间引导路由(Domain-Spatial-Guided Routing,DSR),用于从空间上下文捕获局部拓扑变化;以及熵控制动态分配(Entropy-controlled Dynamic Allocation,EDA),用于根据路由不确定性调整激活专家的数量。这些分支结合了稳定的跨领域表征学习与自适应专家分配。在涵盖室内和室外场景的各类任务上开展的大量实验证明了STAR的有效性:其在ScanNet验证集上达到80.1%的平均交并比(mIoU),在S3DIS上达到77.2%的mIoU,始终优于强基线方法。代码可在我们的项目页面获取(该URL)。

英文摘要

Constructing a unified 3D scene understanding model has long been hindered by the topological discrepancies across sensor modalities. While applying the Mixture-of-Experts (MoE) architecture is a flexible approach for multi-domain 3D understanding, we observe that conventional feature-only MoE routers may underrepresent local sampling topology under semantic supervision, making expert allocation difficult when semantic consistency coexists with geometric heterogeneity. To overcome this challenge, we propose STAR (Spatial-Topology Aware Routing Framework). Specifically, we introduce a multi-attribute self-supervised pre-training branch, covering topological and textural variations, to anchor cross-domain structural priors. Building upon this, we design a domain-aware expert branch with two mechanisms: Domain-Spatial-Guided Routing (DSR), which captures local topological variations from spatial context, and Entropy-controlled Dynamic Allocation (EDA), which adjusts the number of activated experts according to routing uncertainty. Together, these branches combine stable cross-domain representation learning with adaptive expert allocation. Extensive experiments across various tasks, encompassing both indoor and outdoor scenes, demonstrate the effectiveness of STAR. It achieves 80.1% mIoU on the ScanNet validation set and 77.2% mIoU on S3DIS, consistently improving over strong baselines. Code is available at our project page (https://xmw666.github.io/STAR/).

CommentsThe third author is the corresponding author

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑