arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越当前场景:基于事件引用的主动视角选择抓取

Beyond the Current Scene: Event-Referential Grasping with Active View Selection

Hyunjoon Lee, Haebeom Jung, Eunsung Cha, Daeun Lee, Yu-Chiang Frank Wang, Jaesung Choe, Jaesik Park

arXiv 2609.39375首次发表:更新:

发表机构

Seoul National University; NVIDIA(首尔大学; 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出BeyondSCe零样本抓取系统,利用事件历史与主动视角选择,在目标被遮挡时仍能完成事件引用抓取,显著提升成功率并减少视角数。

AI 中文摘要

一个观察人与物体交互的机器人应当能够执行后来引用这些交互的请求。此类请求可能通过目标在过去事件中扮演的角色来指定抓取目标,而非通过其名称或外观。此外,当机器人被要求行动时,目标可能已不再可见。我们提出了BeyondSCe,一个用于这种事件引用场景的零样本机器人抓取系统。给定事件历史和当前场景,系统识别所请求的物体或部件,并定位其以进行抓取。如果目标被遮挡,系统将历史中恢复的事件先验与当前场景几何相结合,选择可能揭示目标的相机视角。该系统使用预训练模型,无需额外的任务特定训练。在带有单个腕装RGB-D相机的真实机器人实验中,对于初始可见和遮挡目标,抓取成功率分别达到76%和77%,而每种条件下最强基线的成功率分别为40%和55%。在四个额外的高遮挡场景中,与给定目标真实3D边界框的主动感知基线相比,抓取成功率从75%提高到95%,同时平均视角数从3.35减少到2.20。

英文摘要

A robot that observes people interacting with objects should be able to carry out later requests that refer back to those interactions. Such requests may specify a grasp target by the role it played in a past event rather than by its name or appearance. Moreover, the target may no longer be visible when the robot is asked to act. We present BeyondSCe, a zero-shot robotic grasping system for this event-referential setting. Given the event history and the current scene, the system identifies the requested object or part and localizes it for grasping. If the target is occluded, it combines an event prior recovered from the history with current scene geometry to select camera viewpoints likely to reveal the target. The system uses pretrained models without additional task-specific training. In real-robot experiments with a single wrist-mounted RGB-D camera, it achieves grasp success rates of 76% and 77% for initially visible and occluded targets, respectively, compared with 40% and 55% for the strongest baseline in each condition. On four additional scenes with heavy occlusion, it increases grasp success rates from 75% to 95% while reducing the mean number of views from 3.35 to 2.20, compared with an active-perception baseline given the target's ground-truth 3D bounding box.

CommentsProject page: https://www.haebeom.com/BeyondCSe/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑