发表机构
Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究双手机器人移动操作中多帧动作去噪问题,提出帧混合策略(MoF),通过在多坐标框架同步去噪,维护规范扩散状态并融合噪声预测,在模拟和实际任务中均优于单帧基线。
AI 中文摘要
机器人操作本质上是多帧的:局部动作在末端执行器框架中可能简单,而运输、直立物体处理和全身协调在基对齐框架中表现更好。然而,现代基于扩散的视觉运动策略通常采用单个预定义动作框架,迫使一个去噪器对该框架中往往不必要复杂的动作分布进行建模。我们提出了帧混合策略(MoF),一种在多个坐标框架上执行同步动作去噪的扩散策略。MoF 维护单个规范扩散状态,在几个与任务相关的框架中重新表达它,应用特定于框架的去噪器,并将它们的噪声预测融合回规范框架。为了使中间噪声扩散状态能够做到这一点,我们在 SE(3)动作参数化中引入了基于列的 6D 旋转表示,该表示支持精确、可微的框架变换,而无需噪声旋转位于 SO(3)流形上。在九个模拟双手机器人操作任务中,我们表明最佳动作框架取决于任务,并且 MoF 优于 oracle 框架选择和标准专家混合(MoE)基线。我们还在两个实际双手机器人移动操作任务上评估了 MoF,证明它优于所有组成的单帧基线。项目主页:这个 https URL
英文摘要
Robotic manipulation is inherently multi-frame: local actions may be simple in an end-effector frame, while transport, upright-object handling, and whole-body coordination are better represented in a base-aligned frame. However, modern diffusion-based visuomotor policies typically commit to a single predefined action frame, forcing one denoiser to model action distributions that are often unnecessarily complex in that frame. We propose Mixture of Frames Policy (MoF), a diffusion policy that performs synchronized action denoising across multiple coordinate frames. MoF maintains a single canonical diffusion state, re-expresses it in several task-relevant frames, applies frame-specialized denoisers, and fuses their noise predictions back in the canonical frame. To make this possible for intermediate noisy diffusion states, we introduce a column-based 6D rotation representation within an SE(3) action parameterization that supports exact, differentiable frame transformations without requiring noisy rotations to lie on the SO(3) manifold. Across nine simulated bimanual manipulation tasks, we show that the best action frame is task-dependent and that MoF improves over oracle frame selection and standard Mixture-of-Experts (MoE) baselines. We further evaluate MoF on two real-world bimanual mobile manipulation tasks, demonstrating that it outperforms all constituent single-frame baselines. Project homepage: https://mofpo.github.io