arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05137cs.CV

SmartMage:面向3D场景理解的动态模态编排

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

发表机构浙江大学 · 斯坦福大学
查看机构详情
  • Zhejiang University(浙江大学)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

Yue Zhang, Yingzhao Jian, Yunqiu Xu, Xiaoxiao Sun, Hehe Fan

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对现有MLLMs采用固定模态组合的缺陷,提出SmartMage模型,通过SMART和MAGE模块动态编排模态,在5个3D场景理解基准上取得最优性能,在RGB视频理解基准上表现具竞争力。

中文摘要 AI 辅助

理解3D场景是具身智能的基础,需要对视觉、几何等多种模态的异构信息进行联合推理。然而,不同查询下各模态的相关性往往存在差异。现有的多模态大语言模型(MLLMs)通常依赖固定的模态组合,忽略了查询依赖的模态需求。这种僵化的设计会引入无关模态的语义噪声,同时未能充分利用信息更丰富的模态,导致计算资源浪费和推理效果被削弱。为应对这些挑战,本文提出SmartMage,一种用于语义感知3D场景理解的统一MLLM,可动态编排异构模态。具体而言,SmartMage包含两个核心模块:(1)语义引导的模态自适应路由(SMART)模块,该模块利用语义先验、文本-模态对齐以及模态质量来选择与任务相关的模态;(2)模态感知门控专家(MAGE)模块,该模块借助模态先验引导专家激活,促进多模态推理中的自适应专业化。实验结果表明,SmartMage在5个3D场景理解基准上达到了当前最优性能,在仅RGB的视频理解基准上也取得了具有竞争力的结果。在我们的诊断基准ScanFacet中,任务被划分为细粒度语义类别,可分析每种语义类型偏好的模态组合,观察到的模态-语义模式进一步验证了SmartMage的有效性。项目页面:this https URL。

英文摘要

Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric cues. However, the relevance of these modalities often varies across queries. Existing Multimodal Large Language Models (MLLMs) typically rely on fixed modality combinations, overlooking query-dependent modality needs. Such a rigid design can introduce semantic noise from irrelevant modalities while underutilizing more informative ones, leading to wasted computation and diluted reasoning. To address these challenges, this paper proposes SmartMage, a unified MLLM that dynamically orchestrates heterogeneous modalities for semantic-aware 3D scene understanding. Specifically, SmartMage incorporates: (1) a Semantic-guided Modality Adaptive RouTing (SMART) module that selects task-relevant modalities using semantic priors, text-modality alignment, and modality quality; and (2) a Modality-Aware Gating Expert (MAGE) module that leverages modality priors to guide expert activation, fostering adaptive specialization in multimodal reasoning. Empirically, SmartMage achieves state-of-the-art performance across five 3D scene understanding benchmarks, and attains competitive results on RGB-only video understanding benchmarks. In our diagnostic benchmark ScanFacet, tasks are divided into fine-grained semantic categories, enabling analysis of modality combinations preferred by each semantic type. The observed modality-semantic patterns provide further evidence of SmartMage's effectiveness. Project page: https://yuecheong.github.io/SmartMage/.

补充信息

↑