基于VLM的目标搜索中的空间-语义不确定性:平衡探索与识别
Spatial-Semantic Uncertainty in VLM-Based Target Search: Balancing Exploration and Identification
浏览论文内容
中文总结 AI 辅助
本研究提出空间-语义不确定性分解方法,通过独立信念和期望信息增益规划,在VLM目标搜索中平衡探索与识别,实验表明该方法能加速决策且不损失准确性。
中文摘要 AI 辅助
机器人根据自然语言描述搜索目标时,不仅需要确定搜索地点,还需要确定哪个观察到的候选对象是期望的目标。这些决策反映了两种不同的不确定性来源——候选位置的空间不确定性和目标身份语义的不确定性——而在基于VLM的搜索系统中,这两种不确定性常常被混为一谈。我们提出了一种空间-语义不确定性公式,该公式对每个组成部分分别维护独立的信念,并将概率性VLM证据整合到全局目标身份后验中,包括未发现目标的概率质量。这种分解使得基于信息论的规划器能够通过空间和语义期望信息增益(EIG)独立地评估候选发现和目标消歧的价值,为在更广泛的探索与更早的识别之间进行权衡提供了显式机制。我们在500个合成目标上评估了六种VLM不确定性引出接口,结果表明相似的识别准确率可能掩盖校准和虚假置信度方面的显著差异。在观测退化的搜索与识别实验中,基于EIG的规划器在75.0%-92.5%的试验中达到自信决策,而随机搜索仅为20.0%,同时不同的空间-语义权重在达到置信度后实现了相当的识别准确率。增加语义重点减少了不必要的探索和VLM查询,表明显式地规划语义不确定性可以加速目标解析而不牺牲决策质量。这些结果突显了不确定性表示和不确定性驱动规划在具身VLM系统中的不同作用。
英文摘要
Robots searching for a target from a natural-language description must determine not only where to search, but also which observed candidate is the desired target. These decisions reflect two distinct sources of uncertainty - spatial uncertainty over candidate locations and semantic uncertainty over target identity - that are often conflated in VLM-based search systems. We introduce a spatial-semantic uncertainty formulation that maintains separate beliefs over each component and integrates probabilistic VLM evidence into a global target-identity posterior, including probability mass for undiscovered targets. This decomposition allows an information-theoretic planner to independently value candidate discovery and target disambiguation through spatial and semantic expected information gain (EIG), providing an explicit mechanism for trading broader exploration against earlier identification. We evaluate six VLM uncertainty-elicitation interfaces on 500 synthetic targets and show that similar recognition accuracy can conceal substantial differences in calibration and false confidence. In degraded-observation search-and-identify experiments, EIG-based planners reach confident decisions in 75.0%-92.5% of trials, compared with 20.0% for Random search, while different spatial-semantic weightings achieve comparable identification accuracy once confidence is attained. Increasing semantic emphasis reduces unnecessary exploration and VLM queries, demonstrating that explicitly planning over semantic uncertainty can accelerate target resolution without sacrificing decision quality. These results highlight the distinct roles of uncertainty representation and uncertainty-driven planning in embodied VLM systems.
发表机构
- Temple University(天普大学)
- University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。