arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2510.11072cs.ROcs.AIcs.LGcs.SYeess.SY

PhysHSI:迈向真实世界可泛化且自然的人形机器人-场景交互系统

PhysHSI: Towards a Real-World Generalizable and Natural Humanoid-Scene Interaction System

Huayi Wang, Wentao Zhang, Runyi Yu, Tao Huang, Junli Ren, Feiyu Jia, Zirui Wang, Xiaojie Niu, Xiao Chen, Jiahe Chen, Qifeng Chen, Jingbo Wang, Jiangmiao Pang

首次发表 更新
浏览论文内容

中文总结 AI 辅助

提出PhysHSI系统,通过仿真中的对抗性运动先验策略学习实现自然泛化动作,结合真实部署中LiDAR与相机的粗到细定位模块,在搬运、坐、躺、站四类任务中取得高成功率和自然运动模式。

中文摘要 AI 辅助

部署人形机器人与真实世界环境交互(例如搬运物体或坐在椅子上)需要可泛化的、逼真的动作以及鲁棒的场景感知。尽管先前的方法分别推进了各项能力,但将它们结合在一个统一系统中仍是一项持续的挑战。在本研究中,我们提出了一种物理世界人形机器人-场景交互系统 PhysHSI,使人形机器人能够自主执行多样化的交互任务,同时保持自然和逼真的行为。PhysHSI 包含一个仿真训练流程和一个真实世界部署系统。在仿真中,我们采用基于对抗性运动先验的策略学习,以在多样化场景中模仿自然的人形机器人-场景交互数据,实现泛化和逼真行为。对于真实世界部署,我们引入了一个由粗到细的物体定位模块,结合 LiDAR 和相机输入,以提供连续且鲁棒的场景感知。我们在仿真和真实世界环境中验证了 PhysHSI 在四个代表性交互任务(搬运箱子、坐下、躺下和站立)上的表现,展示了持续的高成功率、跨多样化任务目标的强泛化能力以及自然的运动模式。

英文摘要

Deploying humanoid robots to interact with real-world environments--such as carrying objects or sitting on chairs--requires generalizable, lifelike motions and robust scene perception. Although prior approaches have advanced each capability individually, combining them in a unified system is still an ongoing challenge. In this work, we present a physical-world humanoid-scene interaction system, PhysHSI, that enables humanoids to autonomously perform diverse interaction tasks while maintaining natural and lifelike behaviors. PhysHSI comprises a simulation training pipeline and a real-world deployment system. In simulation, we adopt adversarial motion prior-based policy learning to imitate natural humanoid-scene interaction data across diverse scenarios, achieving both generalization and lifelike behaviors. For real-world deployment, we introduce a coarse-to-fine object localization module that combines LiDAR and camera inputs to provide continuous and robust scene perception. We validate PhysHSI on four representative interactive tasks--box carrying, sitting, lying, and standing up--in both simulation and real-world settings, demonstrating consistently high success rates, strong generalization across diverse task goals, and natural motion patterns.

补充信息

↑