arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21878cs.CV

ViSMoE:面向具身指代表达接地的视觉感知稀疏混合专家模型

ViSMoE: Visual-Aware Sparse Mixture-of-Experts for Embodied Referring Expression Grounding

Shuo Feng, Piji Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有具身指代表达接地方法无法区分视角与物体导致表示模糊的问题,提出带视觉感知路由策略的ViSMoE框架,在REVERIE和SOON数据集上性能优于现有SOTA方法。

中文摘要 AI 辅助

具身指代表达接地是指使智能体能够在真实环境中导航并根据自然语言指令定位远处物体的任务。在该场景中,智能体每一步需要选择一个视角进行导航,并在目的地的所有候选物体中识别特定物体。然而,大多数现有方法无法区分视角与物体,而是使用普通视觉编码器对其进行处理,导致视角和物体的表示模糊。为解决上述问题,我们提出ViSMoE,它为具身智能体配备了带有视觉感知路由策略的稀疏混合专家(Sparse Mixture-of-Experts)框架。该框架专门处理不同类型的视觉信息,从而生成视角和物体的判别性视觉表示。在REVERIE和SOON数据集上的实验结果表明,ViSMoE的性能优于之前的最先进方法,显示了所提方法的优越性。

英文摘要

Embodied Referring Expression Grounding is the task of enabling an agent to navigate in real environments and to localize a remote object based on natural language instructions. In this scenario, the agent needs to select one view for navigation at each step and identify a specific object among all candidate objects at the destination. However, most of the previous approaches fail to distinguish between views and objects, instead processing them using the vanilla vision encoder, which results in ambiguous representations of both views and objects. To address the above issues, we propose ViSMoE, which equips sparse Mixture-of-Experts with a visual-aware routing policy for the embodied agent. This framework processes different types of visual information specifically, resulting in discriminative visual representations for both views and objects. Experimental results on REVERIE and SOON datasets demonstrate that ViSMoE outperforms the previous state-of-the-art methods, showing the superiority of our proposed method.

发表机构

  • College of Artificial Intelligence(人工智能学院)
  • Nanjing University of Aeronautics and Astronautics(南京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑