弥合学习视觉感知与符号信念空间规划之间的鸿沟
Bridging Learned Visual Perception and Symbolic Belief-Space Planning
- Technion - Israel Institute of Technology(以色列理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出VLM作为概率接地器的新范式,将视觉-语言模型的谓词接地不确定性建模为符号状态上的概率分布,实现信念空间规划,在模拟家庭机器人环境中提升了鲁棒性和任务成功率。
AI中文摘要:
在部分可观测环境中,智能体必须在缺乏完整世界状态知识的情况下行动,并依赖不确定的状态估计流程。在这种不确定性下获得有根据且可验证的符号规划仍然是一个关键挑战。近期工作已集成视觉-语言模型(VLMs)来弥合感知与符号推理之间的鸿沟,遵循两种主要范式。第一种是VLM作为规划器,直接将图像映射为动作序列;第二种是VLM作为接地器,将观测结果接地为符号谓词,作为现成规划器的初始状态。这两种方法都忽略了规划过程中的不确定性,从而损害了鲁棒性。我们引入了第三种范式,即VLM作为概率接地器,这是一种新颖的方法,将VLM谓词接地的不确定性捕获为符号状态上的概率分布。这使得在信念空间中进行规划成为可能,并在不确定性下产生鲁棒的规划。在模拟家庭机器人环境中的实验表明,与确定性接地相比,我们的方法提高了鲁棒性和任务成功率,突显了我们的方法如何利用基础模型在不确定性下实现可靠的规划。
英文摘要:
In partially observable settings, agents must act without full knowledge of the world state and rely on uncertain state-estimation pipelines. Obtaining grounded and verifiable symbolic plans under such uncertainty remains a key challenge. Recent work has integrated Vision-Language Models (VLMs) to bridge perception and symbolic reasoning, following two main paradigms. The first, VLM-as-planner, maps images directly to action sequences, and the second, VLM-as-grounder, grounds observations into symbolic predicates used as the initial state by off-the-shelf planners. Both approaches ignore uncertainty in the planning process, compromising robustness. We introduce a third paradigm, VLM-as-probabilistic-grounder, a novel approach that captures the uncertainty of VLM predicate groundings as a probability distribution over symbolic states. This enables planning in belief space and producing robust plans under uncertainty. Experiments in simulated household robot settings show improved robustness and task success over deterministic grounding, underscoring how our approach leverages foundation models for reliable planning under uncertainty.