AI 中文总结
针对全景图像在情节记忆具身问答中的几何畸变与背景干扰问题,提出基于立方体贴图投影、BLIP-2相关性估计及多样性贪婪选择的视角选择方法,在OpenEQA上取得最先进性能。
AI 中文摘要
具身问答(EQA)要求智能体根据视觉观察回答关于周围环境的自然语言问题。在这项工作中,我们聚焦于开放词汇的情节记忆具身问答(EM-EQA),其中智能体利用记录的观察历史来回答自由形式的问题。全景图像对此任务很有前景,因为它们提供宽视场观察,无需显式相机旋转即可捕获周围上下文。然而,全景图像给EQA带来了两个挑战:(i)等距柱状投影导致严重的几何畸变,降低了视觉-语言模型(VLM)的识别准确性;(ii)将等距柱状图像直接输入VLM会引入过多的无关背景信息,降低答案准确性并增加视觉标记负担。为应对这些挑战,我们提出了一种用于EM-EQA的全景图像视角选择方法。我们的方法通过立方体贴图投影将等距柱状观察转换为透视视图,利用微调后的BLIP-2估计问题条件下的相关性,并通过多样性感知的贪婪选择来选择信息丰富且多样的视角。在OpenEQA的Habitat-Matterport 3D(HM3D)子集上的实验表明,在报告的等距柱状观察模型结果中,我们的方法达到了最先进的模型性能。此外,在移除旋转视图(将观察帧减少65.5%)后,我们的方法在很大程度上保持了其答案准确性。
英文摘要
Embodied Question Answering (EQA) requires agents to answer natural language questions about surrounding environments from visual observations. In this work, we focus on open-vocabulary episodic-memory EQA (EM-EQA), where an agent answers free-form questions using recorded observation histories. Omnidirectional images are promising for this task, as they provide wide field-of-view observations that can capture surrounding context without requiring explicit camera rotations. However, omnidirectional images introduce two challenges for EQA: (i) equirectangular projection causes severe geometric distortion that degrades vision-language model (VLM) recognition accuracy, and (ii) feeding equirectangular images directly into VLMs introduces excessive irrelevant background information, reducing answer accuracy and increasing the visual-token burden. To address these challenges, we propose a viewpoint selection method for EM-EQA using omnidirectional images. Our method converts equirectangular observations into perspective views via cubemap projection, estimates question-conditioned relevance with fine-tuned BLIP-2, and selects informative and diverse viewpoints through diversity-aware greedy selection. Experiments on the Habitat-Matterport 3D (HM3D) subset of OpenEQA show that our method achieves state-of-the-art model performance among the reported model results with equirectangular observations. Moreover, after removing rotation views, which reduces observation frames by 65.5%, our method largely maintains its answer accuracy.
CommentsAccepted to ACCV 2026. Supplementary video is provided as an ancillary file