发表机构
Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对杂乱场景中任务导向抓取时可用性区域被遮挡的问题,提出ATAP框架,通过生成式形状先验想象完整表面并预测可用性分布,结合不确定性感知的视角规划器迭代验证,显著提升功能抓取成功率并减少57%以上的NBV步骤。
AI 中文摘要
任务导向抓取(TOG)要求机器人抓取物体的功能部件(例如,为倒水而抓取杯子的手柄),然而在杂乱场景中这些可用性区域经常被遮挡。通过下一最佳视角(NBV)规划进行的主动感知可以通过移动相机获取更多信息性观察来解决此类遮挡问题。然而,现有的NBV方法通常优化视角以整体抓取目标物体,而不区分哪个部件与任务相关。一种朴素的适应方法是,在预测可用性之前完全扫描目标物体,这会将大部分视角预算浪费在与任务无关的表面上(例如,为倒水而扫描杯体)。为解决此问题,我们提出了ATAP,一个可用性目标主动感知框架,将视角规划从详尽的目标扫描转变为针对性的可用性验证。ATAP通过生成式形状先验假设被遮挡的目标几何形状,并在想象出的完整表面上预测可用性分布。在杂乱场景中,严重遮挡可能使隐藏可用性的位置变得模糊,鉴于部分观察,多个位置都可能成立。因此,ATAP引入了一个不确定性感知的视角规划器,该规划器联合优化这些竞争假设的预期熵减少和来自真实观察的预期可用性验证增益。此过程迭代进行,直到可用性被充分验证以执行抓取。在仿真和真实世界杂乱场景中的实验表明,ATAP在功能抓取成功率上显著优于固定视角的TOG基线,并且相比基于重建的主动感知,NBV步骤减少了超过57%。
英文摘要
Task-oriented grasping (TOG) requires robots to grasp functional parts of objects (e.g., the handle of a mug for pouring), yet these affordance regions are frequently occluded in cluttered scenes. Active perception via next-best-view (NBV) planning can resolve such occlusions by moving the camera for more informative observations. However, existing NBV methods typically optimize viewpoints for grasping the target object as a whole without distinguishing which part is task-relevant. A naive adaptation, fully scanning the target object before predicting the affordance, wastes most of the viewpoint budget on task-irrelevant surfaces (e.g., the mug body for pouring). To address this, we propose ATAP, an Affordance-Targeted Active Perception framework that shifts viewpoint planning from exhaustive target scanning to targeted affordance verification. ATAP hypothesizes the occluded target geometry via a generative shape prior and predicts the affordance distribution over the imagined complete surface. In cluttered scenes, severe occlusion can make the location of the hidden affordance ambiguous, leaving multiple locations plausible given the partial observation. ATAP therefore introduces an uncertainty-aware viewpoint planner that jointly optimizes expected entropy reduction over these competing hypotheses and expected affordance verification gain from real observations. This process iterates until the affordance is sufficiently verified for grasp execution. Experiments in simulation and real-world cluttered scenes show that ATAP substantially improves the functional grasp success rate over fixed-view TOG baselines, and outperforms reconstruction-based active perception with over 57% fewer NBV steps.