RynnBrain 1.1:迈向更具能力和通用性的具身基础模型
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
浏览论文内容
中文总结 AI 辅助
介绍RynnBrain 1.1具身基础模型家族,其经统一框架训练,相比1.0版本有改进。开发RynnBrain - VLA并部署。该模型在多方面成果显著,初始化策略表现优,联合训练提升了得分和成功率。
中文摘要 AI 辅助
我们展示了RynnBrain 1.1,这是一个涵盖2B、9B和122B - A10B规模的具身基础模型家族。它通过统一的时空和物理基础框架进行训练,支持具身感知、空间推理、定位和规划。与RynnBrain 1.0相比,它在整个模型家族中进一步引入了接触点预测,并为2B和9B模型引入了原生3D基础。我们还开发了具有统一跨具身动作空间和特定具身掩码的RynnBrain - VLA,并将其部署在Unitree G1、Astribot - S1和Tianji - Wuji上。RynnBrain 1.1在具身认知、定位和3D基础方面取得了显著成果,122B - A10B模型在VSI - Bench、MMSI和RefSpatial - Bench上优于所有评估的专有和开源模型。真实机器人实验表明,以RynnBrain初始化的策略优于基于Qwen的和有代表性的通用VLA,而联合多任务和多具身训练比单任务训练提高了过程得分和成功率。
英文摘要
We present RynnBrain 1.1, a family of embodied foundation models spanning 2B, 9B, and 122B-A10B scales. Trained with a unified spatio-temporal and physically grounded framework, RynnBrain 1.1 supports embodied perception, spatial reasoning, localization, and planning. Compared with RynnBrain 1.0, it further introduces contact-point prediction across the model family and native 3D grounding for the 2B and 9B models, yielding representations and outputs that are more directly aligned with robot manipulation. We also develop RynnBrain-VLA with a unified cross-embodiment action space and embodiment-specific masking, and deploy it on Unitree G1, Astribot-S1, and Tianji-Wuji. RynnBrain 1.1 achieves strong results on embodied cognition, localization, and 3D grounding, with the 122B-A10B model outperforming all evaluated proprietary and open-source models on VSI-Bench, MMSI, and RefSpatial-Bench. Real-robot experiments show that RynnBrain-initialized policies outperform Qwen-based and representative generalist VLAs, while joint multi-task and multi-embodiment training improves process scores and success rates over per-task training.
发表机构
- DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团)
- Hupan Lab(湖畔实验室)
机构由 AI 辅助整理,请以论文原文为准。