FlipToSee:通过重新抓取进行主动视觉探索的概率稳定放置先验
FlipToSee: A Probabilistic Stable Placement Prior for Active Visual Exploration via Regrasping
查看机构详情
- King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
FlipToSee提出一种基于单视角点云的概率放置先验,通过参数化支撑法线并利用von Mises-Fisher混合密度网络,实现多模态稳定放置预测,在仿真和物理机器人上显著提升主动视觉探索的重新抓取成功率。
中文摘要 AI 辅助
对桌面物体的主动视觉探索通常需要将未知的静止物体重新定向到不同的稳定支撑面,以暴露被遮挡的表面。为了无需穷举物理搜索即可识别此类放置,我们从单视角点云中学习一个概率放置先验。稳定放置预测本质上是多模态的,而传统的6自由度回归通过建模平移和平面内偏航引入了进一步的歧义。因此,我们提出FlipToSee,一个概率框架,通过将放置参数化为$S^2$上的单位支撑法线来消除这种表示歧义,同时通过von Mises--Fisher混合密度网络建模其多模态条件分布。为了将模式多样性与物理鲁棒性解耦,FlipToSee从混合分量中确定性地提取一个紧凑的候选集,并使用通过候选对齐监督训练的辅助头进行鲁棒性感知的重新排序。在仿真中,FlipToSee在分布内物体上实现了$98.4\%$的首次提议成功率,在分布外形状上实现了$95.3\%$,在零样本迁移到家用YCB物体上实现了$90.0\%$。我们还通过将学习到的放置先验与抓取和运动规划集成,在物理机器人上进行了探索性重新抓取的演示。
英文摘要
Active visual exploration of tabletop objects often requires reorienting an unknown resting object onto a different stable support face to expose occluded surfaces. To identify such placements without exhaustive physical search, we learn a probabilistic placement prior from a single-view point cloud. Stable placement prediction is inherently multimodal, and conventional 6-DoF regression introduces further ambiguity by modeling translation and in-plane yaw. We therefore propose FlipToSee, a probabilistic framework that removes this representational ambiguity by parameterizing placements as unit support normals on $S^2$ while modeling their multimodal conditional distribution via a von Mises--Fisher mixture density network. To decouple mode diversity from physical robustness, FlipToSee deterministically extracts a compact candidate set from the mixture components and applies robustness-aware reranking using an auxiliary head trained with candidate-aligned supervision. In simulation, FlipToSee achieves $98.4\%$ first-proposal success on in-distribution objects, $95.3\%$ on out-of-distribution shapes, and $90.0\%$ under zero-shot transfer to household YCB objects. We further demonstrate the learned placement prior on a physical robot by integrating it with grasp and motion planning for exploratory regrasping.