发表机构
Princeton University; Toyota Research Institute; Physical Intelligence(普林斯顿大学; 丰田研究所; Physical Intelligence)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EgoLAP提出一种VLA预训练框架,通过语言动作思维链联合学习人类与机器人轨迹,利用运动意图跨具身迁移,在真实世界任务中达80.1%进展,性能提升2.3倍。
AI 中文摘要
自我中心人类数据为扩展机器人学习提供了一条超越昂贵机器人演示的路径,然而具身差距使得原始人类轨迹难以成为控制的有效监督目标。我们的关键洞见是,尽管低级动作具有具身特异性,但其潜在的运动意图能够捕捉到跨人类和机器人迁移的任务相关结构。我们提出了EgoLAP,一个视觉-语言-动作(VLA)预训练框架,通过共享的基于语言的动作思维链,联合从人类和机器人轨迹中学习。EgoLAP将运动意图表达为结构化的、时间上抽象的语言动作,并将其与基于场景几何、物理和物体可供性的运动级推理相结合。在广泛的真实世界和模拟实验中,EgoLAP将人类经验迁移到机器人控制的效果优于其他动作表示,达到了80.1%的平均真实世界任务进展,相比其他动作表示取得了2.3倍的性能提升。运动级推理也优于结合子任务、物体框和视觉轨迹推理的复合推理格式。
英文摘要
Egocentric human data offer a path to scaling robot learning beyond costly robot demonstrations, yet the embodiment gap makes raw human trajectories a poor supervisory target for control. Our key insight is that, although low-level actions are embodiment-specific, their underlying motion intent can capture task-relevant structure that transfers across humans and robots. We introduce EgoLAP, a VLA pre-training framework that jointly learns from human and robot trajectories through a shared language-based action chain-of-thought. EgoLAP expresses motion intent as structured, temporally abstracted language actions and pairs them with motion-level reasoning grounded in scene geometry, physics, and object affordances. Across extensive real-world and simulated experiments, EgoLAP transfers human experience to robot control more effectively than alternative action representations and reaches 80.1% mean real-world task progress, a 2.3x performance gain over alternative action representations. Motion-level reasoning also outperforms a composite reasoning format that combines subtask, object-box, and visual-trace reasoning.
CommentsProject website: https://ego-lap.github.io/