发表机构
Northeastern University; Stanford University; University of Pennsylvania; Brown University; Texas A&M University(东北大学; 斯坦福大学; 宾夕法尼亚大学; 布朗大学; 德州农工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对将操纵动作表示为相机平面二维轨迹带来的挑战,提出Pix2Act模仿学习方法,通过生成图像空间关键点轨迹及三角测量恢复末端执行器姿态,将三维控制转为二维预测问题,提升泛化能力与性能,优于现有基线且抗相机扰动。
AI 中文摘要
将操纵动作表示为相机平面中的二维轨迹,为学习复杂的三维操纵策略提供了紧凑且可解释的基础。然而,它也带来了帧外轨迹和精度有限的挑战。我们提出了Pix2Act,一种模仿学习方法,通过在每个相机平面中生成连续的图像空间关键点轨迹,并通过三角测量无损地恢复末端执行器姿态来应对这些挑战。这将高维三维控制重新表述为一个更简单、更易学习的二维预测问题。关键的是,它在同一坐标空间中对齐观察和动作,使等变变换能够将单个相机图像与其图像空间动作一起联合旋转。我们分析了这种增强的对称特性,并设计了一种能够融合多个相机视图同时尊重其每个视图旋转的网络架构。结果,Pix2Act隐含地扩大了数据分布的支持范围,并学习了跨变换的不变动作结构,从而提高了泛化能力和整体性能。在各种模拟和现实世界的操纵任务中,Pix2Act优于现有基线,并且在相机扰动下保持稳健。
英文摘要
Representing manipulation actions as 2D trajectories in the camera plane provides a compact and interpretable basis for learning complex 3D manipulation policies. However, it also creates challenges from out-of-frame trajectories and limited precision. We propose Pix2Act, an imitation learning method that addresses these challenges by generating continuous image-space keypoint trajectories in each camera plane and losslessly recovering end-effector poses via triangulation. This reformulates high-dimensional 3D control as a simpler, more learnable 2D prediction problem. Crucially, it aligns observations and actions in the same coordinate space, enabling equivariant transformations to jointly rotate individual camera images together with their image-space actions. We analyze the symmetry properties of this augmentation and design a network architecture that can fuse multiple camera views while respecting their per-view rotations. As a result, Pix2Act implicitly enlarges the support of the data distribution and learns invariant action structures across transformations, yielding improved generalization and overall performance. Across diverse simulated and real-world manipulation tasks, Pix2Act outperforms state-of-the-art baselines and remains robust under camera perturbations.
CommentsProject Website: https://haojhuang.github.io/pix2act_page/