发表机构
Robotic Systems Lab, ETH Zurich(苏黎世联邦理工学院机器人系统实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出GENESIS-Handover方法,利用VLM图像生成任务特定的手-物体交互假设,实时匹配人类手部姿态,实现稳健的任务导向机器人-人交接,用户研究中83.3%认为其任务理解优于先前方法。
AI 中文摘要
当人类互相传递物体时,他们会将几何和语义信息都融入这一过程。例如,将刀以手柄朝向接收者的方式传递,而非刀刃朝向,既更符合人体工程学也更安全。近期最先进的任务导向机器人-人交接方法已从建模物体几何进展到融入物体可供性。然而,它们往往忽略了预测人类为使用物体而选择的具体、任务特定的手部姿态。由于许多物体支持多种交互模式,例如羊角锤可用于敲击或拔钉,这种变异性必须被建模以实现稳健的任务导向交接。为解决这一问题,我们提出了一种新颖方法,GENESIS-Handover(生成式假设选择),它利用VLM图像生成来产生多种任务特定的手-物体交互假设。这些假设被实时与观察到的人类手部姿态进行匹配,从而推断出最合适的交接配置。通过利用VLM作为合理手-物体交互的先验,该方法为未见过的物体-任务对生成任务条件下的交接策略。我们在将完整系统部署到移动操作器上之前,先评估了独立的交互提议模块。在一项涉及12名参与者、跨越五个任务-物体对的用户研究中,83.3%的参与者认为我们的方法比先前最先进的方法具有更好的任务理解能力。
英文摘要
When humans hand each other objects, they incorporate both geometric and semantic information into this process. For example, passing a knife with the handle towards the recipient, rather than the blade, is both more ergonomic and safer. Recent state-of-the-art methods for task-oriented robot-human handovers have progressed from modeling object geometry to incorporating object affordances. However, they often forgo predicting the explicit, task-specific hand poses a human selects to utilize an object. Since many objects support multiple interaction modalities, e.g., a claw hammer used to strike or pull nails, this variability must be modeled to achieve robust task-oriented handovers. To tackle this, we propose a novel approach, GENESIS-Handover (GENErative HypotheSIS), which leverages VLM image generation to produce a variety of task-specific hand-object interaction hypotheses. These hypotheses are matched in real time to the observed human hand pose, enabling inference of the most suitable handover configuration. By leveraging VLMs as priors of plausible hand-object interactions, the method produces task-conditioned handover strategies for previously unseen object-task pairs. We evaluate the standalone interaction proposal module before deploying the full system on a mobile manipulator. In a user study with 12 participants across five task-object pairs, 83.3% perceived our method to have better task understanding than the previous state of the art.
CommentsAccepted to the Conference on Robot Learning (CoRL) 2026