发表机构
Tencent Robotics X; Hy Vision Team; Futian Laboratory(腾讯Robotics X团队; 腾讯混元视觉团队; 福田实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在构建物理世界具身智能体,介绍Hy-Embodied-VLM-1.0模型。定义以行动为中心的能力分类法,开发数据管道。基于特定主干和编码器构建模型,用专家混合架构提升效率。在多基准测试中性能出色,较上一代有显著提升,在具身智能任务中也表现强大。
AI 中文摘要
构建有能力的具身智能体不仅需要多模态感知和理解,还需要行动推理、适应不断变化的情况以及与物理世界交互的智能能力。在本报告中,我们介绍了Hy-Embodied-VLM-1.0,这是一个专门为在物理世界中运行的具身智能体设计的高效且强大的具身基础模型。从预训练阶段开始培养这些能力,我们定义了一个以行动为中心的能力分类法,包括三个递进维度:与行动相关的状态理解、行动转换推理以及顺序和自适应推理。在此分类法指导下,我们开发了系统的数据管道并策划了涵盖预训练和训练后的数据混合。为了在支持对延迟敏感的部署的同时提供强大的物理世界理解和交互能力,我们基于Hy3-A3B语言主干和Hy-ViT2视觉编码器构建模型。其高效的专家混合架构将强大的模型能力与高推理效率结合在一起。我们在一套涵盖具身感知、物理世界理解和具身推理的38个基准测试中对Hy-Embodied-VLM-1.0进行了评估。该模型在38个基准测试中的19个上在同等规模模型中取得了最佳性能,并且大幅超越了强大的竞争对手,包括Qwen3.6-A3B和Cosmos 3。与上一代Hy-Embodied-0.5 MoT-2B相比,Hy-Embodied-VLM-1.0将平均性能提高了8.4%。尽管仅激活了3B参数,但它实现了与激活32B参数的上一代模型相近的性能。除了静态基准测试评估之外,Hy-Embodied-VLM-1.0在需要多轮交互和长视野推理的具身智能任务上也表现出强大的性能。
英文摘要
Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, adapting to evolving situations, and interacting with the physical world. In this report, we introduce Hy-Embodied-VLM-1.0, an efficient and powerful embodied foundation model specifically designed for embodied agents operating in the physical world. To cultivate such capabilities from the pre-training stage onward, we define an action-centric capability taxonomy comprising three progressive dimensions: Action-Relevant State Understanding, Action-Transition Reasoning, and Sequential and Adaptive Reasoning. Guided by this taxonomy, we develop a systematic data pipeline and curate data mixtures spanning both pre-training and post-training. To deliver strong physical-world understanding and interaction capabilities while supporting latency-sensitive deployment, we build our model on the Hy3-A3B language backbone and the Hy-ViT2 vision encoder. Its efficient Mixture-of-Experts architecture combines strong model capacity with high inference efficiency. We evaluate Hy-Embodied-VLM-1.0 on a comprehensive suite of 38 benchmarks covering embodied perception, physical-world understanding, and embodied reasoning. The model achieves the best performance among similarly sized models on 19 of the 38 benchmarks and substantially outperforms strong competitors, including Qwen3.6-A3B and Cosmos 3. Compared with the previous-generation Hy-Embodied-0.5 MoT-2B, Hy-Embodied-VLM-1.0 improves average performance by 8.4%. Despite activating only 3B parameters, it achieves performance close to that of the previous-generation model with 32B activated parameters. Beyond static benchmark evaluation, Hy-Embodied-VLM-1.0 also demonstrates strong performance on embodied agentic tasks requiring multi-turn interaction and long-horizon reasoning.
CommentsTech Report. Code and models are open-sourced at https://github.com/Tencent-Hunyuan/HY-Embodied