arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05674cs.RO

JoyAI-RA 0.5:通过双动作对齐扩展机器人操纵学习

JoyAI-RA 0.5: Scaling Robot Manipulation Learning via Dual Action Alignment

RA Team

首次发表
浏览论文内容

中文总结 AI 辅助

针对机器人操纵学习中异构数据的负迁移问题,提出 JoyAI-RA 0.5 框架,通过双动作对齐结合 VLWA 与强化学习,在 AgiBot 基准上实现性能提升,证明人类弱标注数据可作为扩展操纵能力的核心资源。

中文摘要 AI 辅助

机器人数据稀缺,因此通用策略需要从异构源学习,包括人类第一视角视频、仿真和真实机器人,这些源在监督和 embodiment(具身性)方面存在差异,存在动作标签缺失或互不兼容的问题。人类第一视角数据规模最大,但与机器人数据的差距最远,直接汇集会导致负迁移而非知识共享。我们提出 JoyAI-RA 0.5,这是一种通用视觉-语言-世界-动作(VLWA)框架,它将物理世界动力学先验与视觉语义相结合,并通过双动作对齐在这类数据上扩展操纵学习。隐式动作对齐从视觉过渡中推断潜在动作,使无动作的人类、仿真和机器人数据能够指导以潜在动作为条件的世界模型学习物理动力学。显式对齐通过规范动作表示和相机帧分块相对末端执行器动作,将可靠的人类和机器人轨迹锚定在统一的物理动作空间中。随后的内外环强化阶段将高效的任务适应与基础策略改进相结合。在真实世界的 AgiBot 基准测试中,JoyAI-RA 在已见任务和未见变体上均表现出色。任务分数随着人类第一视角预训练数据量的增加而持续提升,在我们的最大规模下未出现 plateau(平台)迹象。这表明,大量但弱标注的人类经验可转化为可迁移的训练信号,使人类视频不再仅仅是弱辅助源,而是操纵能力可沿其扩展的主要轴。项目页面可在该 https URL 找到。

英文摘要

Robot data is scarce, so generalist policies need to learn from heterogeneous sources, including human egocentric video, simulation, and real robots, which differ in supervision and embodiment, with action labels missing or mutually incompatible. Human egocentric data scale best but sit farthest from robot data, and naive pooling causes negative transfer rather than knowledge sharing. We propose JoyAI-RA 0.5, a generalist Vision-Language-World-Action (VLWA) framework that couples physical world-dynamics priors with visual semantics and scales manipulation learning across such data via dual action alignment. Implicit action alignment infers latent actions from visual transitions, enabling action-free human, simulation, and robot data to guide a latent-action-conditioned world model in learning physical dynamics. Explicit alignment grounds reliable human and robot trajectories in a unified physical action space through a canonical action representation and camera-frame chunk-relative end-effector actions. An inner-outer-loop reinforcement stage then pairs efficient task adaptation with foundation-policy improvement. On a real-world AgiBot benchmark, JoyAI-RA performs strongly on both seen tasks and unseen variations. The task score improves consistently as the volume of human egocentric pretraining data increases and shows no sign of plateauing at our largest scale. This suggests that abundant but weakly labeled human experience can be converted into a transferable training signal, making human video not merely a weak auxiliary source but a primary axis along which manipulation capability can be scaled. Project page can be found at https://joyai-ra-05.github.io/.

发表机构

  • Joy Future Academy, JD(京东探索研究院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑