arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14187cs.AIcs.RO

RxBrain:具有联合语言-视觉推理与想象的具身认知基础模型

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination

Haotian Liang, Mingkang Chen, Yufei Huang, Yuchun Guo, Xiaomeng Zhu, Xiangli Shi, Kaixuan Wang, Yunxuan Mao, Weijie Zhou, Ling Chen, Shirong Zeng, Yueyu Long, Y… 展开作者

Haotian Liang, Mingkang Chen, Yufei Huang, Yuchun Guo, Xiaomeng Zhu, Xiangli Shi, Kaixuan Wang, Yunxuan Mao, Weijie Zhou, Ling Chen, Shirong Zeng, Yueyu Long, Yuchen Si, Yajuan Zhu, Xingyu Zhou, Minghui Wang, Wanjia He, Xin Yang, Lingzhu Xiang, Zhiqing Liu, Bohan Ma, Xiran Huang, Tianshuo Yang, Zhiheng Liu, Xuantang Xiong, Zisheng Lu, Ping Luo, Yao Mu, Han Hu, Zhengyou Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

研究提出具身认知基础模型RxBrain,采用统一多模态架构,通过构建自动管道训练,引入RxBrain-Bench评估,能保持具身理解和生成能力,扩展到连续机器人动作生成,为具身认知基础模型发展奠定基础。

中文摘要 AI 辅助

具身认知要求智能体将高级任务推理与要实现的物理状态联系起来。我们引入了Hy-Embodied-RxBrain,一个具有联合语言-视觉推理与想象的具身认知基础模型。与强调场景理解和文本决策的视觉语言模型或主要预测未来视觉状态的生成世界模型不同,RxBrain在单个规划序列中表示具身计划,语言和视觉想象发挥互补作用。语言提供计划的抽象结构,视觉想象通过世界状态预测和联合子目标规划来支撑该结构。RxBrain采用统一的多模态Transformer混合架构。为训练此能力,构建了自动管道将具身视频转换为联合文本-视觉规划监督。还引入RxBrain-Bench评估模型。实验表明RxBrain保持具身理解和生成能力,还扩展到连续机器人动作生成,为具身认知基础模型迈出了第一步。

英文摘要

Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoning and imagination. Unlike vision-language models that emphasize scene understanding and textual decision making, or generative world models that mainly predict future visual states, RxBrain represents embodied plans in a single planning sequence where language and visual imagination play complementary roles. Language provides the abstract structure of a plan, including task decomposition, planning primitives, constraints, temporal order, and decision logic, while visual imagination grounds this structure through world state prediction and joint subgoal planning, associating each planning step with intermediate and final physical states. RxBrain adopts a unified multimodal Mixture-of-Transformers architecture that supports language, image, and video understanding and generation within one model. To train this capability, we build an automatic pipeline that converts embodied videos into joint text-visual planning supervision by decomposing videos into planning steps and aligning them with visual state transitions. We further introduce RxBrain-Bench to evaluate whether models can represent embodied plans through joint textual and visual components rather than separate understanding or generation. Experiments show that RxBrain maintains embodied understanding and generation abilities, and produces plans with coupled textual reasoning, world state prediction, and joint subgoal planning. We also extend RxBrain to continuous robot action generation, where it shows promising real-robot performance without large-scale action-data pretraining. These results provide an initial step toward foundation models for embodied cognition.

↑