arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AcrossWAM1.0:用于紧凑机器人策略的模块化潜在世界-动作栈

AcrossWAM1.0:A Modular Latent World-Action Stack for Compact Robot Policies

Yafei Zhang, Nan Wu

arXiv 2608.29937首次发表:更新:

发表机构

Across Physical AI; Institute of Automation, Chinese Academy of Sciences(跨物理人工智能机构; 中国科学院自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对LaWAM的耦合问题推出AcrossWAM1.0模块化潜在世界-动作栈,在LIBERO实验中实现高成功率,参数规模显著缩小,提供可审计的部署边界。

AI 中文摘要

潜在世界-动作模型通过在特征空间中预测与动作相关的视觉子目标,避免了未来像素的渲染。LaWAM确立了这一框架,但其原始呈现将世界模型、多模态主干和部署检查点紧密耦合。我们推出AcrossWAM1.0,这是对该潜在世界-动作栈的模块化与扩展性研究。我们未将潜在子目标作为新算法提出,而是明确了模块边界:策略适配器生成潜在动作和动作生成上下文;保留的潜在世界解码器将预测的转换锚定在当前场景中;流匹配专家生成连续动作块。我们进一步将仅用于训练的教师与推理图分离,并提供可验证的部署导出。在2000个配对的LIBERO episode上,将Qwen3-VL-2B主干替换为Qwen3.5-0.8B后,成功率为97.45%,而2B模型的成功率为98.00%(差值为0.55个百分点;精确McNemar检验p=0.266)。这并未证明二者等效,但符合预先设定的2点保留标准。该紧凑、可推理的检查点包含1472.6M个唯一参数,比原始2B策略少42.4%,且所有保留的张量与源检查点位对位相同。我们还通过MiniCPM-V适配器冒烟测试检查了跨系列执行,闭环跨系列迁移仍为待评估问题。因此,AcrossWAM1.0为紧凑的潜在世界-动作策略提供了可审计的软件和评估边界,这与LaWAM原始的潜在子目标贡献有所区别。

英文摘要

Latent world-action models avoid rendering future pixels by predicting an action-relevant visual subgoal in feature space. LaWAM established this formulation, but its original presentation left the world model, multimodal backbone, and deployment checkpoint tightly coupled. We introduce AcrossWAM1.0, a modularization and scaling study of this latent world-action stack. Rather than presenting latent subgoals as a new algorithm, we make the module boundary explicit: a policy adapter produces latent-action and action-generation contexts; a retained latent world decoder grounds the predicted transition in the current scene;and a flow-matching expert generates continuous action chunks. We further separate training-only teachers from the inference graph and provide a verifiable deployment export. On 2,000 paired LIBERO episodes, replacing a Qwen3-VL-2B backbone with Qwen3.5-0.8B yields 97.45% success versus 98.00% for the 2B model (a-0.55percentage-point difference; exact McNemarp=0.266). This does not prove equivalence, but it meets a prespecified two-point retention criterion. The compact, inference-reachable checkpoint contains 1,472.6M unique parameters, 42.4% fewer than the original 2B policy, while all retained tensors are bitwise identical to the source checkpoint. Cross-family execution is additionally checked with a MiniCPM-V adapter smoke test; closed-loop cross-family transfer remains an open evaluation. AcrossWAM1.0 therefore contributes an auditable software and evaluation boundary for compact latent world-action policies, distinct from LaWAM's original latent-subgoal contribution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑