arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从像素到位姿:基于人类演示的以物体为中心的工具操作学习

From Pixel to Poses: Object-centric Tool Manipulation Learning from Human Demonstrations

Bangjun Wang, Longyan Wu, Yukun Wei, Shenghe Shao, Chaoyi Huang, Wenze Cui, Zetong Xu, Hanlin Wu, Long Chen, Yi Ma, Hongyang Li

arXiv 2609.35375首次发表:更新:

发表机构

The University of Hong Kong; Shanghai Innovation Institute; Xiaomi EV(香港大学; 上海创新研究院; 小米电动汽车)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对机器人操作数据稀缺及复杂工具操作精度不足的问题,提出P2P-T框架,通过两阶段方法从人类演示中学习,无需人机对齐数据,实现73%的性能提升。

AI 中文摘要

扩展机器人操作能力的主要瓶颈在于真实世界机器人数据的稀缺。虽然近期的方法利用人类视频演示来缓解这一短缺,但这些方法计算成本高昂,并且仍然依赖配对的人机数据进行领域对齐。尽管当前最先进的方法在长时程任务中表现出色,但在复杂工具操作所需的精细和精确控制方面仍存在困难。为克服这些限制,我们提出了P2P-T(从像素到位姿的工具操作,Pixel to Poses for Tool Manipulation),一个数据高效、以物体为中心的框架,可直接从人类演示中学习工具使用。P2P-T通过两阶段方法弥合认知与物理执行之间的鸿沟:首先,预训练一个以物体为中心的世界模型以提取稳定的位姿先验;其次,将这些先验集成到一个高效、位姿感知的低层策略中。通过利用由现代基础模型驱动的稳健自动化数据处理流水线,P2P-T完全绕过了对人机对齐数据的需求,从而大幅降低了整体训练开销。在最小的每任务微调下,我们的框架在复杂的真实世界工具操作任务上,执行性能相比先前最先进方法提升了73%,而这些任务目前对于标准的大规模预训练模型来说仍难以企及。

英文摘要

Scaling up robotic manipulation is primarily bottlenecked by the scarcity of real-world robot data. While recent approaches leverage human video demonstrations to mitigate this shortage, they remain computationally expensive and still rely on paired human-robot data for domain alignment. Although current state-of-the-arts excel at long-horizon tasks, they struggle with the delicate and precise control required for complex tool manipulation. To overcome these limitations, we introduce P2P-T, from Pixel to Poses for Tool Manipulation, a data-efficient, object-centric framework that learns tool use directly from human demonstrations. P2P-T bridges the cognitive and physical execution gap through a two-stage approach. First, pretraining an object-centric world model to extract stable pose priors; second, integrating these priors into an efficient, pose-aware low-level policy. By utilizing a robust automated data processing pipeline powered by modern foundation models, P2P-T completely bypasses the need for human-robot aligned data. This reduces overall training overhead drastically. With minimal per-task fine-tuning, our framework achieves a 73% improvement over the previous state of the art in execution performance on complex, real-world tool manipulation tasks that currently remain out of reach for standard large-scale pretrained models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑