发表机构
Microsoft Research(微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Rho是一个开源权重VLA模型家族,通过动作专家架构和本体中间训练,在三种双臂机器人上实现高效数据轻量适配,并具备基于纠正反馈的在线适配能力,性能优于现有模型。
AI 中文摘要
通用物理AI模型必须将广泛的视觉和语言能力与跨机器人本体的精确控制以及下游任务的高效适配相结合。我们引入了Rho,一个面向双臂操作的开源权重VLA模型家族,专为在代表研究实验室和工业界双臂机器人的3种本体上进行数据轻量级任务适配而设计——YAM Box、UR AI Trainer和FR3 Duo。我们系统地消融了Rho的动作专家架构和训练方案,并在受控仿真和物理机器人实验中表明,本体中间训练可改善下游适配。由此产生的针对YAM Box、UR AI Trainer和FR3 Duo的Rho变体匹配或超越了现有的开源权重VLA,并在本报告评估的任务、本体和基线中取得了最强的整体性能。我们进一步展示了Rho模型家族内置的在线适配能力:一个内部潜在策略从纠正性反馈中学习,为冻结的流匹配动作专家选择观测条件噪声输入。仅需少至15个纠正情节,适配这一轻量模块即可使Rho处理其离线微调分布边缘的任务情境。这些结果共同将Rho定位为既是一个强大的通用机器人操作模型,又是一个实用的适配基础。我们发布了基础Rho模型和特定本体检查点,以促进Rho在研究实验和实际工业用例中的部署。
英文摘要
General-purpose physical AI models must combine broad visual and linguistic capabilities with precise control across robot embodiments and efficient adaptation to downstream tasks. We introduce Rho, a family of open-weights VLA models for bimanual manipulation designed for data-light task adaptation on 3 embodiments representative of dual-arm robots across research labs and the industry -- YAM Box, UR AI Trainer, and FR3 Duo. We systematically ablate Rho's action-expert architecture and training recipe, and show in controlled simulation and physical-robot experiments that embodiment midtraining improves downstream adaptation. The resulting Rho variants for YAM Box, UR AI Trainer, and FR3 Duo match or outperform existing open-weights VLAs and achieve the strongest overall performance across the tasks, embodiments, and baselines evaluated in this report. We further demonstrate the Rho model family's built-in capacity for online adaptation: an internal latent policy learns from corrective feedback to select observation-conditioned noise inputs for the frozen flow-matching action expert. With as few as 15 corrected episodes, adapting this lightweight module enables Rho to handle task situations at the fringe of its offline finetuning distribution. Together, these results position Rho as both a strong general-purpose robotic manipulation model and a practical foundation for adaptation. We release the base Rho model and the embodiment-specific checkpoints to facilitate Rho's deployment in research experiments and practical industrial use cases.