EWAM:统一具身模型中的涌现深度专业化——从语义理解经由视觉预见再到行动
EWAM: Emergent Depth-Wise Specialization in a Unified Embodied Model -- From Semantic Understanding through Visual Foresight to Action
- Yinwang Intelligent Technology Co. Ltd(银旺智能科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
EWAM提出一种以行动为中心的统一具身模型,通过非对称联合注意力在无逐层监督下涌现深度专业化,在仿真和真实机器人上超越现有基线,并证明统一学习可诱导从语义理解到行动形成的内部有序演进。
AI中文摘要:
视觉-语言-行动(VLA)策略强调语义理解,而世界-行动模型(WAM)则学习环境动力学的预测性表征。将策略同时暴露于这两种来源的系统,往往仍将行动计算集中于单一专家模型上。我们提出EWAM,一种以行动为中心的统一具身模型,其非对称联合注意力机制使得行动令牌能够在每一层读取语义、当前视觉、预测未来以及行动信息,同时感知专家保持其各自不同的角色。在没有逐层监督的情况下,EWAM展现出一种涌现的深度专业化:行动查询在浅层主要关注视觉-语言特征,在中间层关注预测的未来帧,而在深层则关注行动令牌本身。这种交接在不同任务间复现,并在去噪步骤中保持稳定。检查点追踪和因果干预表明,这种专业化是习得的,且行动生成依赖于它。EWAM在两种独立的预训练机制下进行训练,一种基于跨具身机器人轨迹,另一种基于人类第一人称视频。在仿真和真实机器人实验中,它超越了现有的VLA、WAM及混合基线模型。人类第一人称数据同时提升了跨具身迁移和真实机器人鲁棒性,而子任务阶段监督则改善了长时程任务的完成率。综合这些结果表明,统一具身学习能够诱导从语义理解、经由视觉预见、再到行动形成的有序内部演进过程。
英文摘要:
Vision-language-action (VLA) policies emphasize semantic understanding, whereas world-action models (WAMs) learn predictive representations of environment dynamics. Systems that expose a policy to both sources often still concentrate action computation on a single expert. We present EWAM, an action-centric unified embodied model whose asymmetric joint attention lets action tokens read semantic, current-visual, predicted-future, and action information at every layer while the perceptual experts retain their distinct roles. Without layer-wise supervision, EWAM develops an emergent depth-wise specialization: action queries attend mainly to vision-language features in shallow layers, to predicted future frames in intermediate layers, and to action tokens themselves in deep layers. This handoff replicates across tasks and is stable across denoising steps. Checkpoint tracking and causal interventions show that it is learned and that action generation depends on it. EWAM is pretrained in two separate regimes, one on cross-embodiment robot trajectories and one on human egocentric video. In simulation and real-robot experiments, it surpasses existing VLA, WAM, and hybrid baselines. Human egocentric data improve both cross-embodiment transfer and real-robot robustness, and subtask-phase supervision improves long-horizon completion. Together, these results suggest that unified embodied learning can induce an ordered internal progression from semantic understanding, through visual foresight, to action formation.