发表机构
Luleå University of Technology(吕勒奥理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对VLM在非结构化环境导航中的语义模糊性问题,提出基于高斯上下文分布的可通行性估计流程,利用概念锚定生成稠密可通行性与不确定性地图,在GOOSE数据集上验证了有效性。
AI 中文摘要
在非结构化环境中的自主导航需要鲁棒的场景理解,然而视觉语言模型(VLMs)常常遭受语义模糊性的困扰,其中相互冲突的预测可能导致危险的故障。为了解决这一问题,我们提出了一种新颖的基于视觉的可通行性估计流程,该流程显式地建模上下文不确定性。我们的方法利用概念锚定(Conceptual Anchoring)将开放词汇的VLM预测锚定到连续的物理可通行性尺度上。通过将模型的响应表述为高斯上下文分布(GCD),我们基于该分布的统计特性推导出稠密的可通行性地图和稠密的不确定性地图。在真实世界的GOOSE数据集上的实验验证表明,我们提出的不确定性度量有效地与模糊性来源(如视觉伪影和混合地形重叠)相关联。该方法在提供统计不确定性估计以解决语义模糊性的独特优势的同时,展现出具有竞争力的性能,从而在复杂的户外环境中实现更安全、更可靠的自主行为。
英文摘要
Autonomous navigation in unstructured environments requires robust scene understanding, yet Vision-Language Models (VLMs) often suffer from semantic ambiguity, where conflicting predictions can lead to dangerous failures. To address this, we present a novel pipeline for vision-based traversability estimation that explicitly models contextual uncertainty. Our approach utilizes Conceptual Anchoring to ground open-vocabulary VLM predictions onto a continuous physical traversability scale. By formulating the model's responses as a Gaussian Context Distribution (GCD), we derive both a dense traversability map and a dense uncertainty map based on the statistical properties of the distribution. Experimental validation on the real-world GOOSE dataset demonstrates that our proposed uncertainty metric effectively correlates with sources of ambiguity, such as visual artifacts and mixed terrain overlap. The method exhibits competitive performance while offering the distinct advantage of providing statistical uncertainty estimates to address semantic ambiguity, enabling safer and more reliable autonomous behavior in complex outdoor settings.