部分可观测条件下语言引导物体检索的选择性承诺
Selective Commitment for Language-Guided Object Retrieval under Partial Observability
- University of Arizona(亚利桑那大学)
- King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对部分可观测下的语言引导物体检索,提出闭环框架,通过联合信念、VLM观测和有限时域规划协调证据收集与选择性抓取承诺,在模拟和真实试验中验证了有效性。
AI中文摘要:
在部分可观测条件下进行语言引导的物体检索,需要决定是收集更多证据、与场景交互、抓取候选物体,还是弃权(不执行)。我们提出一个闭环框架,用于协调这些决策,以检索与参考容器相关联的目标物体。该框架通过跟踪物体、未观测目标和目标缺失三种假设,维护关于目标身份、容器关系和存在性的持续联合信念。视图条件下的分类视觉语言模型观测更新该信念;保形抓取资格和机器人可行性控制承诺,而有限时域信念空间规划选择信息收集动作。在五种不同场景中,我们提出的方法在19/25次模拟试验中成功,而最佳的任务适应基线为12/25,并且是唯一在每个场景中至少取得一次成功的评估策略。消融研究表明,跨视图记忆在部分遮挡下提高了成功率,而完整系统并未始终优于简化变体。真实机器人试验展示了闭环重新观测和从注入的抓取失败中自主恢复,而注入的视角失败导致错误的弃权(不执行)。实验结果证明了在部分可观测条件下,在统一框架内协调证据收集和选择性抓取承诺的可行性。
英文摘要:
Language-guided object retrieval under partial observability requires deciding whether to gather more evidence, interact with the scene, grasp a candidate, or abstain. We present a closed-loop framework that coordinates these decisions for retrieving a target specified in relation to a reference container. The framework maintains a persistent joint belief over target identity, container relation, and presence through tracked-object, unobserved-target, and target-absent hypotheses. View-conditioned categorical VLM observations update this belief; conformal grasp eligibility and robot feasibility govern commitment, while finite-horizon belief-space planning selects information-gathering actions. Across five different scenarios, our proposed method succeeds in 19/25 simulation episodes versus 12/25 for the best-performing task-adapted baseline and is the only evaluated policy to achieve at least one success in each scenario. Ablations show that cross-view memory improves success under partial occlusion, while the full system does not consistently outperform simplified variants. Real-robot trials demonstrate closed-loop re-observation and autonomous recovery from injected grasp failures, while injected viewpoint failures end in false defer. Experimental results demonstrate the feasibility of coordinating evidence gathering and selective grasp commitment within a unified framework for retrieval under partial observability.