arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Praxis:从自我中心视频中提炼物理交互先验,实现全身操作的可泛化性

Praxis: Distilling Physical Interaction Priors from Egocentric Videos for Generalizable Whole-Body Manipulation

Shuliang He, Ruiyan Xu, Bo Yue, Hengming Zhang, Huayi Zhou, Shuai Wang, Wei-Shi Zheng, Guiliang Liu

arXiv 2609.30735首次发表:更新:

AI 中文总结

Praxis通过从一次性自我中心视频中提炼物理交互先验,结合闭环姿态校准和在线感知,实现移动人形全身操作的空间、视觉及跨物体泛化,并抵抗外部干扰。

AI 中文摘要

移动人形操作要求既达到可用的工作空间,又要在物体姿态和接触条件变化时保持精确的手-物交互。从有限的任务特定数据中学习这些行为仍然具有挑战性。为弥合这一差距,我们引入了Praxis,一个全身操作框架,该框架结合了来自一次性自我中心视频演示的物理交互先验、闭环姿态校准和在线感知。该框架协调三个阶段:视觉-语言引导的朝向目标物体的导航、闭环姿态校准以对齐手臂-手工作空间,以及同步上下身控制的灵巧操作。在线视觉反馈在新物体姿态和场景配置下重新锚定演示的交互几何,而触觉反馈则使手部运动适应实际接触条件。每个操作技能由一个人类演示指定,无需任务特定的操作策略重新训练。在五个长时程操作任务上的实验展示了空间、视觉和跨物体泛化,以及所有三个阶段中对外部物理干扰的恢复能力。

英文摘要

Mobile humanoid manipulation requires both reaching a usable workspace and preserving precise hand-object interactions as object poses and contact conditions change. Learning these behaviors from limited task-specific data remains challenging. To bridge this gap, we introduce Praxis, a whole-body manipulation framework that combines physical interaction priors from one-shot egocentric video demonstrations with closed-loop posture calibration and online perception. The framework coordinates three stages: vision-language-guided navigation toward target objects, closed-loop posture calibration to align the arm-hand workspace, and dexterous manipulation with synchronized upper- and lower-body control. Online visual feedback re-grounds demonstrated interaction geometry under new object poses and scene configurations, while tactile feedback adapts hand motions to actual contact conditions. Each manipulation skill is specified by one human demonstration, without task-specific manipulation-policy retraining. Experiments across five long-horizon manipulation tasks demonstrate spatial, visual, and cross-object generalization, as well as recovery from external physical disturbances across all three stages.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑