IMPRINT:用于长尾目标导航的图像条件查询增强
IMPRINT: Image-Conditioned Query Enrichment for Long-Tail Object Goal Navigation
浏览论文内容
中文总结 AI 辅助
研究针对具身人工智能中ObjectNav依赖纯文本查询可靠性低的问题,提出IMPRINT框架,用网络图像丰富文本查询改善定位,无需训练导航策略。通过新基准HSSD-rare评估,显示图像条件查询可提升导航增益,还指出下游检测质量是关键瓶颈。
中文摘要 AI 辅助
具身人工智能越来越依赖于由预训练视觉语言模型构建的可查询语义地图,以实现零样本目标导航(ObjectNav)。然而,现有方法通常依赖于纯文本查询,随着语义特异性向细粒度对象类别增加,其可靠性会降低。我们引入了IMPRINT,这是一个零样本即插即用框架,通过网络源图像丰富文本对象查询,以改善在可查询地图中的定位。检索到的图像使用视觉语言模型进行编码,与语义地图匹配以生成相似性地图,并进行聚合以产生上下文感知定位。值得注意的是,这不需要对基础导航策略进行训练或修改。为了明确评估长尾行为,我们提出了HSSD-rare,这是一个基于Habitat合成场景构建的新ObjectNav基准,具有语义特定的子类别。在OVON和HSSD-rare中,图像条件查询都能持续改善对象定位并产生端到端导航增益。进一步分析表明,将定位增益转化为导航性能关键取决于下游检测质量,突出了长尾具身导航中的一个关键系统瓶颈。
英文摘要
Embodied AI increasingly relies on queryable semantic maps built from pre-trained vision-language models to enable zero-shot Object Goal Navigation (ObjectNav). However, existing approaches typically depend on text-only queries, which become less reliable as semantic specificity increases toward fine-grained object categories. We introduce IMPRINT, a zero-shot plug-and-play framework that enriches textual object queries with web-sourced images to improve grounding in queryable maps. Retrieved images are encoded using a vision-language model, matched against the semantic map to produce similarity maps, and aggregated to yield context-aware localization. Notably, this requires no training or modification of the underlying navigation policy. To explicitly evaluate long-tail behavior, we present HSSD-rare, a new ObjectNav benchmark built on Habitat Synthetic Scenes and featuring semantically specific subcategories. Across both OVON and HSSD-rare, image-conditioned queries consistently improve object grounding and yield end-to-end navigation gains. Further analysis reveals that translating localization gains to navigation performance depends critically on downstream detection quality, highlighting a key systems bottleneck in long-tail embodied navigation.
发表机构
- University of Padova(帕多瓦大学)
- Fondazione Bruno Kessler (FBK)(布鲁诺·凯斯勒基金会)
机构由 AI 辅助整理,请以论文原文为准。