arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24526cs.CV

ME-VLM:用于具身认知与智能体协调的统一视觉语言模型

ME-VLM: A Unified VLM for Embodied Cognition and Agent Coordination

  • Li Auto Inc(理想汽车)

机构由 AI 辅助整理,请以论文原文为准。

Foundation Model, Li Auto Inc

AI总结:

ME-VLM是一个统一的视觉语言模型,通过结合具身认知与多模态智能体能力,采用多教师蒸馏训练,在具身和智能体基准上表现优异,并支持边缘设备高效推理。

AI中文摘要:

物理人工智能要求模型在真实世界环境中将视觉和语言理解落地,同时考虑环境约束和执行反馈。我们提出了MachEmbodied-VLM(ME-VLM),一个统一的视觉语言模型,具有4B和35B-A3B两个变体,将具身认知和多模态智能体能力集于一身。我们的工作强调物理感知和时空推理,以及在数字和物理环境中的规划、交互和结果评估。我们构建了涵盖具身和多模态智能体任务的训练数据,包括执行观察和反馈,以支持结果评估和决策改进。训练流程包括具身能力注入、具身和多模态智能体专家的分别强化学习,以及多教师在线策略蒸馏,将它们的互补能力整合到单一模型中。实验表明,在具身和智能体基准测试以及自动驾驶和具身导航任务上均具有竞争力的性能。对于边缘部署,视觉令牌压缩、W4A8量化和软硬件协同优化使得4B变体在M100上实现设备端推理,将预填充延迟从400毫秒降低到188毫秒。项目页面:此https链接 代码仓库:此https链接

英文摘要:

Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints and execution feedback. We introduce MachEmbodied-VLM (ME-VLM), a unified vision-language model with two variants, 4B and 35B-A3B, that brings together embodied cognition and multimodal agent capabilities. Our work emphasizes physical perception and spatiotemporal reasoning, together with planning, interaction, and outcome assessment in both digital and physical environments. We construct training data spanning embodied and multimodal agent tasks, including execution observations and feedback to support outcome assessment and decision refinement. The training pipeline comprises embodied capability injection, separate reinforcement learning of embodied and multimodal-agent experts, and multi-teacher on-policy distillation that consolidates their complementary capabilities into a single model. Experiments show competitive performance on both embodied and agent benchmarks, as well as on autonomous-driving and embodied-navigation tasks. For edge deployment, visual token compression, W4A8 quantization, and hardware-software co-optimization enable on-device inference of the 4B variant on the M100, reducing prefill latency from 400 ms to 188 ms. Project Page: https://machembodied.com/ME-Brain/ME-VLM.html Code Repository: https://github.com/MachEmbodied/ME-VLM

↑