arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18718cs.RO

面向杂乱环境抓取的基于视觉语言模型的校准概率性障碍推理

Calibrated Probabilistic Obstruction Reasoning with Vision-Language Models for Grasping in Clutter

  • University of Engineering and Technology, Vietnam National University, Hanoi(越南国立大学河内工程技术大学)
  • Japan Advanced Institute of Science and Technology(日本先端科学技术大学院大学)
  • Hanyang University(汉阳大学)

机构由 AI 辅助整理,请以论文原文为准。

Thanh-Tuan Tran, Ngoc-Chien Chu, Thanh Nguyen Canh, Nak Young Chong, Nguyen-Viet Ha, Xiem HoangVan

中文总结 AI 辅助

提出CPOR-Grasp框架,通过校准和融合VLM、深度与amodal线索,对障碍图进行概率推理与边缘化,实现可认证的抓取决策,显著降低校准误差并提升成功率。

中文摘要 AI 辅助

从杂乱环境中取回目标物体需要决定是抓取目标、移除障碍物还是弃权(不执行)。现有方法通常致力于单一的障碍图或移除策略,忽略了不同场景解释之间的不确定性。它们还依赖于校准不佳的视觉语言模型(VLM)预测,并可能产生联合不一致的成对障碍关系。此外,当前的近似方法无法保证丢弃的假设对最终决策的影响。我们提出了CPOR-Grasp,一种校准的概率性障碍推理框架,将成对证据的不确定性传播到行动决策中。CPOR-Grasp校准并融合VLM、深度和amodal掩码线索来估计障碍概率,在有效障碍图上诱导分布,并对这些图进行边缘化以计算目标可访问或应移除给定障碍物的可能性。为了使推理易于处理,它仅保留最高概率的图,并推导出丢弃概率质量的总变差界,从而实现认证决策、自适应停止和原则性弃权(不执行)。在合成和真实的UNOBench场景中,CPOR-Grasp优于最先进的基线。校准误差在Gemini Robotics骨干上从0.1416降至0.0185,而图截断在99.74%的决策中与精确推理匹配,使用的图数量减少了56倍。在真实世界实验中,CPOR-Grasp实现了77.8%的平均成功率,超过了SOTA基线。

英文摘要

Retrieving a target from clutter requires deciding whether to grasp the target, remove a blocker, or defer. Existing methods typically commit to a single obstruction graph or removal strategy, ignoring uncertainty across alternative scene interpretations. They also rely on miscalibrated vision-language model (VLM) predictions and can produce pairwise obstruction relations that are jointly inconsistent. Moreover, current approximations provide no guarantees about the impact of discarded hypotheses on the final decision. We propose CPOR-Grasp, a calibrated probabilistic obstruction-reasoning framework that propagates uncertainty from pairwise evidence to action decisions. CPOR-Grasp calibrates and fuses VLM, depth, and amodal-mask cues to estimate obstruction probabilities, induces a distribution over valid obstruction graphs, and marginalizes over these graphs to compute the likelihood that the target is accessible or that a given blocker should be removed. To make inference tractable, it retains only the highest-probability graphs and derives a total-variation bound on the discarded probability mass, enabling certified decisions, adaptive stopping, and principled deferral. On synthetic and real UNOBench scenes, CPOR-Grasp outperforms state-of-the-art baselines. Calibration error decreases from 0.1416 to 0.0185 on the Gemini Robotics backbone, while graph truncation matches exact inference on 99.74\% of decisions using 56 times fewer graphs. In real-world experiments, CPOR-Grasp achieves a 77.8\% average success rate, surpassing SOTA baselines.

补充信息

↑