DeViGrasp:感知退化下四足操作器的鲁棒视觉移动抓取
DeViGrasp: Robust Visual Mobile Grasping for Quadruped Manipulators under Degraded Perception
- Tsinghua Shenzhen International Graduate School, Tsinghua University(清华大学深圳国际研究生院)
- Shanghai Jiao Tong University Global College, Shanghai Jiao Tong University(上海交通大学全球学院)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对四足操作器在视觉感知退化下移动抓取鲁棒性不足的问题,提出DeViGrasp-Bench基准和DeViGrasp-Net教师-学生框架,融合状态条件抓取推理与可靠性感知时间估计,在困难设置下成功率62.3%,显著优于现有方法。
中文摘要 AI 辅助
四足操作器能够在复杂环境中实现移动抓取,然而其全身控制策略仍易受不可靠的机载视觉感知影响。现有方法通常在相对可靠的观测条件下开发,尚未系统性地研究遮挡、分割掩码缺失、深度噪声和目标定位抖动对抓取推理及目标跟踪的影响。为填补这一空白,我们提出了DeViGrasp-Bench,一个面向退化视觉下移动抓取的基准,它包含受控的视觉退化、已知和未知物体、多个难度级别以及复杂地形,并评估任务成功率、执行效率和动作平滑度。我们进一步提出DeViGrasp-Net,一个教师-学生框架,将状态条件抓取推理与可靠性感知的时间目标估计相结合。特权教师网络基于物体、机器人、末端执行器和任务状态,关注离线抓取候选;可部署的学生网络通过目标保持记忆和时间记忆注意力,将双视角分割深度观测与当前、记忆和恢复目标假设融合。DeViGrasp-Net在各级退化、未知物体和复杂地形上均优于VBC,并在所有评估的退化级别上超越改进的DQ-Net。在困难设置下,其成功率达到62.3%,分别比VBC和DQ-Net提高16.1和4.3个百分点;在较难设置下,其相对于DQ-Net的优势增至10.5个百分点。消融研究证实了抓取感知监督与可靠性感知时间记忆的互补优势。
英文摘要
Quadruped manipulators enable mobile grasping in complex environments, yet their whole-body control policies remain vulnerable to unreliable onboard visual perception. Existing methods are typically developed under relatively reliable observations and have not systematically examined how occlusion, segmentation-mask dropout, depth noise, and target-localization jitter affect grasp reasoning and target tracking. To address this gap, we introduce DeViGrasp-Bench, a benchmark for mobile grasping under degraded vision that incorporates controlled visual degradations, seen and unseen objects, multiple difficulty levels, and complex terrains, and evaluates task success, execution efficiency, and action smoothness. We further propose DeViGrasp-Net, a teacher--student framework that combines state-conditioned grasp reasoning with reliability-aware temporal target estimation. The privileged teacher attends to offline grasp candidates conditioned on object, robot, end-effector, and task states, while the deployable student fuses dual-view segmented-depth observations with current, memory, and recovery target hypotheses through Target Hold Memory and Temporal Memory Attention. DeViGrasp-Net outperforms VBC across degradation levels, unseen objects, and complex terrains, and surpasses an adapted DQ-Net across all evaluated degradation levels. Under the Difficult setting, it achieves a success rate of 62.3\%, improving upon VBC and DQ-Net by 16.1 and 4.3 percentage points, respectively; under the Hard setting, its margin over DQ-Net increases to 10.5 percentage points. Ablation studies confirm the complementary benefits of grasp-aware supervision and reliability-aware temporal memory.