发表机构
Xiaopeng(小鹏汽车)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出以执行为中心的具身视觉语言模型Capek 0.5,通过分类法整合四类具身能力,在多维度评估中表现优于初始化模型,可迁移至闭环具身任务执行。
AI 中文摘要
视觉语言模型正日益成为具身智能体的推理核心。机器人的执行本质上是迭代的:每一个动作都会重塑场景和物理状态,持续更新需要感知、推理和验证的内容。要满足这些需求,需要具备在监督信号、预测格式和验证标准上有所不同的互补能力。现有方法通常针对孤立的、特定任务的目标开发这些能力,留下了如何围绕整体执行对它们进行组织和集成的问题。我们提出了Capek 0.5,这是一个围绕以执行为中心的能力分类法构建的具身视觉语言模型。该分类法不是按数据集或任务组织训练,而是根据具身能力在整个执行过程中的功能角色对其进行分组,包含四个能力家族:空间推理、时间理解、动作指导和状态验证。每个能力首先由一个专用专家通过强化学习获取,使用来自共享主干的可验证奖励,然后通过权重空间合并和路由策略空间蒸馏将这些专家整合为单个推理时模型。我们在2B和35B-A3B规模上实例化了Capek 0.5,并从三个互补的角度对其进行评估:包括Capek-StateBench(一个新的状态验证基准)在内的综合基准套件;从专家到统一模型的能力保留的受控研究;以及模拟具身环境中的闭环评估。Capek 0.5在绝大多数匹配的基准行上优于其初始化模型,在一个检查点中保留了全部四个专业能力并具有量化损失,且可迁移至闭环具身任务执行。
英文摘要
Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewing what must be perceived, reasoned about, and verified. Meeting these demands requires complementary capabilities that differ in supervision signals, prediction formats, and verification criteria. Existing approaches typically develop these capabilities against isolated, task-specific objectives, leaving open how they should be organized and integrated around execution as a whole. We present Capek 0.5, an embodied vision-language model built around an execution-centric capability taxonomy. Rather than organizing training by datasets or tasks, the taxonomy groups embodied capabilities according to their functional roles throughout execution and comprises four capability families: Spatial Reasoning, Temporal Understanding, Action Guidance, and State Verification. Each capability is first acquired by a dedicated specialist through reinforcement learning with verifiable rewards from a shared backbone, and the specialists are then consolidated into a single inference-time model through weight-space merging followed by routed policy-space distillation. We instantiate Capek 0.5 at the 2B and 35B-A3B scales and evaluate it from three complementary perspectives: comprehensive benchmark suites including Capek-StateBench, a new benchmark for state verification; a controlled study of capability retention from specialists to the unified model; and closed-loop evaluation in simulated embodied environments. Capek 0.5 improves the large majority of matched benchmark rows over its initialization, retains all four specialized capabilities in one checkpoint with quantified losses, and transfers to closed-loop embodied task execution.