发表机构
Northeastern University; Stanford University; Brown University(东北大学; 斯坦福大学; 布朗大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对机器人学习中动作空间的挑战,提出动作映射策略(AMP),将3D闭环操作策略学习转为图像空间分类问题,通过投影动作到图像平面,以像素为类别控制维度,实现高精度、快速推理,实验证明其优于基线。
AI 中文摘要
动作空间在机器人学习中是一项重大挑战,因其维度高、时间跨度长且常有多模态最优解。动作表示和损失函数的选择虽能解决部分问题,但存在权衡。我们提出动作映射策略(AMP)将3D闭环操作策略学习转化为图像空间的分类问题。虽分类在生成语言模型中有效,但应用于机器人动作学习困难,因简单离散高维连续动作会使令牌词汇量激增。我们的关键想法是将3D动作投影到相机图像平面,把每个像素位置视为离散类别,控制维度并保留多模态。该方法支持高维动作毫米级精度,无需过大词汇量,保留细粒度像素视觉信号,单步前向就能预测整个动作块,避免复杂噪声调度和迭代去噪,推理速度比扩散策略快得多。在各种操作任务上的实验表明,AMP优于强大基线,成功率更高、推理更快且空间推理能力更强。
英文摘要
The action space poses a major challenge in robot learning, since it is often high-dimensional, can span long time horizons, and frequently admits multi-modal optimal solutions. A good choice of action representation and loss function can help to address these concerns, but there are often trade offs. We propose Action Map Policy (AMP), which casts 3D closed-loop manipulation policy learning as a classification problem in image space. While classification has been an effective formulation in generative language models, applying it to robot action learning is difficult because naively discretizing high-dimensional continuous actions explodes the token vocabulary. Our key idea is to project 3D actions onto the camera image planes and treat each pixel location as a discrete class, thus controlling dimensionality while retaining multi-modality. This method supports millimeter-level precision for high-dimensional actions without requiring a prohibitively large vocabulary, while preserving fine-grained pixel-wise visual signals. Furthermore, it can predict the entire action chunk in a single forward pass, avoiding complex noise scheduling and iterative denoising while achieving substantially faster inference than diffusion policies. Experiments on various manipulation tasks show that AMP outperforms strong baselines, achieving higher success rates, faster inference, and enhanced spatial reasoning.
CommentsProject Website: https://haojhuang.github.io/amp_page/