发表机构
The University of Hong Kong; Shanghai Innovation Institute; Xiaomi EV(香港大学; 上海创新研究院; 小米电动汽车)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机器人操作数据稀缺及复杂工具操作精度不足的问题,提出P2P-T框架,通过两阶段方法从人类演示中学习,无需人机对齐数据,实现73%的性能提升。
AI 中文摘要
扩展机器人操作能力的主要瓶颈在于真实世界机器人数据的稀缺。虽然近期的方法利用人类视频演示来缓解这一短缺,但这些方法计算成本高昂,并且仍然依赖配对的人机数据进行领域对齐。尽管当前最先进的方法在长时程任务中表现出色,但在复杂工具操作所需的精细和精确控制方面仍存在困难。为克服这些限制,我们提出了P2P-T(从像素到位姿的工具操作,Pixel to Poses for Tool Manipulation),一个数据高效、以物体为中心的框架,可直接从人类演示中学习工具使用。P2P-T通过两阶段方法弥合认知与物理执行之间的鸿沟:首先,预训练一个以物体为中心的世界模型以提取稳定的位姿先验;其次,将这些先验集成到一个高效、位姿感知的低层策略中。通过利用由现代基础模型驱动的稳健自动化数据处理流水线,P2P-T完全绕过了对人机对齐数据的需求,从而大幅降低了整体训练开销。在最小的每任务微调下,我们的框架在复杂的真实世界工具操作任务上,执行性能相比先前最先进方法提升了73%,而这些任务目前对于标准的大规模预训练模型来说仍难以企及。
英文摘要
Scaling up robotic manipulation is primarily bottlenecked by the scarcity of real-world robot data. While recent approaches leverage human video demonstrations to mitigate this shortage, they remain computationally expensive and still rely on paired human-robot data for domain alignment. Although current state-of-the-arts excel at long-horizon tasks, they struggle with the delicate and precise control required for complex tool manipulation. To overcome these limitations, we introduce P2P-T, from Pixel to Poses for Tool Manipulation, a data-efficient, object-centric framework that learns tool use directly from human demonstrations. P2P-T bridges the cognitive and physical execution gap through a two-stage approach. First, pretraining an object-centric world model to extract stable pose priors; second, integrating these priors into an efficient, pose-aware low-level policy. By utilizing a robust automated data processing pipeline powered by modern foundation models, P2P-T completely bypasses the need for human-robot aligned data. This reduces overall training overhead drastically. With minimal per-task fine-tuning, our framework achieves a 73% improvement over the previous state of the art in execution performance on complex, real-world tool manipulation tasks that currently remain out of reach for standard large-scale pretrained models.