arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

JAMB:用于双臂操作的动作-运动联合扩散

JAMB: Joint Action-Motion Diffusion for Bimanual Manipulation

Chuyang Xiao, Peilin Meng, David Held

arXiv 2609.25322首次发表:更新:

发表机构

Carnegie Mellon University; University of Michigan(卡内基梅隆大学; 密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对双臂操作中动作与未来几何后果缺乏联合建模的问题,提出JAMB扩散策略,联合去噪动作与未来3D点轨迹,在模拟和真实任务中显著提升成功率与泛化能力。

AI 中文摘要

协调的双臂操作具有挑战性,因为任一臂的运动都可能改变共享的3D场景,从而影响另一臂。然而,大多数扩散策略在生成动作时并未显式建模这些未来的几何后果,而预测变体通常仅将未来状态用作辅助监督或固定条件。我们通过提出JAMB来解决这一局限性,这是一种联合去噪双臂动作和未来3D点轨迹的扩散策略。通过允许动作和轨迹假设在共享的Transformer内共同演化,两者可以在去噪过程中相互告知并细化。我们进一步将多模态表示锚定在共享的时空坐标系中,以促进联合去噪期间的几何感知交互。我们在RoboTwin 2.0中的多样化双臂操作任务以及真实世界机器人上评估了JAMB,将其与仅动作策略以及涵盖不同状态表示和学习目标的替代未来预测方法进行了比较。在16个模拟任务中,JAMB实现了83.4%的平均成功率,比最强基线高出23.9个百分点。在三个真实世界任务中,它分别比仅动作方法和辅助几何预测方法高出50.0和21.2个百分点。除了这些性能提升外,JAMB在杂乱场景和分布外背景上的泛化能力也强于所评估的基线。这些结果共同证明了我们用于协调双臂操作的联合动作-运动建模框架的有效性。我们的项目网站可在以下https URL访问。

英文摘要

Coordinated bimanual manipulation is challenging because the motion of either arm can alter the shared 3D scene and thereby affect the other arm. Yet most diffusion policies generate actions without explicitly modeling these future geometric consequences, while predictive variants typically use future state only as auxiliary supervision or fixed conditioning. We address this limitation by proposing JAMB, a diffusion policy that jointly denoises bimanual actions and future 3D point tracks. By allowing action and track hypotheses to evolve together within a shared Transformer, each can inform and refine the other throughout denoising. We further ground multimodal representations in a shared spatiotemporal coordinate system to facilitate geometry-aware interaction during joint denoising. We evaluate JAMB on diverse bimanual manipulation tasks in RoboTwin 2.0 and on a real-world robot, comparing it with action-only policies and alternative future-prediction approaches spanning different state representations and learning objectives. Across 16 simulation tasks, JAMB achieves an average success rate of 83.4%, outperforming the strongest baseline by 23.9 percentage points. On three real-world tasks, it outperforms the action-only and auxiliary geometry prediction methods by 50.0 and 21.2 percentage points, respectively. Beyond these performance gains, JAMB shows stronger generalization to cluttered scenes and out-of-distribution backgrounds than the evaluated baselines. Together, these results demonstrate the effectiveness of our joint action-motion modeling framework for coordinated bimanual manipulation. Our project website is available at https://jam-bimanual.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑