发表机构
Nanyang Technological University; College of Computing and Data Science; Zhejiang University; College of Computer Science and Technology; Alibaba Group(南洋理工大学; 计算与数据科学学院; 浙江大学; 计算机科学与技术学院; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对人机物体交接数据集稀缺及现实-模拟差距问题,提出Hand2Bot数据集与PassGen生成流水线,实现高意图识别准确率与低误触发率,支持机器人稳健零样本迁移与早期意图预测。
AI 中文摘要
人机(H2R)物体交接是人机协作的基础能力,但大规模以人为中心的数据集稀缺以及显著的现实-模拟差距阻碍了该领域的进展。为应对这些挑战,我们推出Hand2Bot,这是一个RGB-D视频数据集,提供身体姿态、面部表情等丰富上下文信息,专门采集自包含现实噪声模式的交接场景。我们还提出PassGen,这是一种生成式流水线,利用稳定视频扩散模型和意图感知时间面部编码器合成逼真的交接序列,同时确保手-物体一致性。为弥合现实-模拟差距,我们实施基于形态学的深度编辑策略,复制物理深度图中的现实传感器噪声。实验评估表明,我们的框架在消融研究和物理机器人平台的现实部署中均实现了高意图识别准确率和低误触发率。我们的结果证实,在PassGen上训练可实现稳健的零样本迁移,且相比传统以手为中心的基线能更早地进行意图预测,有效在共享工作空间中实现具有社会意识的机器人行为。
英文摘要
Human-to-robot (H2R) object handover is a fundamental capability for human-robot collaboration, yet progress is hindered by the scarcity of large-scale, human-centric datasets and the significant sim-to-real gap. To address these challenges, we introduce Hand2Bot, an RGB-D video dataset that provides rich contextual information such as body posture and facial expressions, specifically collected for handover scenarios with real-world noise patterns. We further propose PassGen, a generative pipeline that leverages stable video diffusion and an Intention-Aware Temporal Face Encoder to synthesize realistic handover sequences while ensuring hand-object consistency. To bridge the sim-to-real gap, we implement a morphology-based depth editing strategy that replicates realistic sensor noise found in physical depth maps. Experimental evaluations demonstrate that our framework achieves high intention identification accuracy and low false trigger rates in both ablation studies and real-world deployment on a physical robot platform. Our results confirm that training on PassGen allows for robust zero-shot transfer and earlier intention anticipation compared to traditional hand-centric baselines, effectively enabling socially aware robotic behavior in shared workspaces.