学习使用想象力:面向世界动作模型的进度条件化未来利用
Learning to Use Imagination: Progress-Conditioned Future Utilization for World Action Models
浏览论文内容
中文总结 AI 辅助
提出ProWAM,一种进度条件化世界动作模型,通过自监督双时间进度编码器和层次化进度条件化想象调制,自适应利用未来想象,提升VLA/WAM模型性能。
中文摘要 AI 辅助
世界动作模型(WAMs)通过将未来视觉动态纳入动作生成,扩展了视觉-语言-动作(VLA)模型。然而,现有的WAMs在利用想象出的未来时,对不断演变的执行进度的适应性有限,可能引入分散注意力或不可靠的预测线索。这一局限性源于经验上识别出的两种未来效用非均匀性:(i)在进度间层面,随着控制需求的变化,想象未来的效用在不同执行阶段有所不同;(ii)在进度内层面,在相同的进度状态下,各个未来潜在表征表现出异质性相关性。为解决这些局限性,我们提出了ProWAM,一种进度条件化世界动作模型,将执行进度作为显式的中间表征引入,以实现自适应的想象力利用。ProWAM包含两个紧密耦合的组件:(1)为获得可靠的执行进度表征,我们提出了自监督双时间进度编码器(SS-DTPE)。SS-DTPE将短期动作-观测交互建模与长期循环进度聚合相结合,以捕捉最近的执行反馈和累积的任务历史。(2)基于SS-DTPE的进度表征,我们提出了层次化进度条件化想象调制(HPIM),使想象力利用适应执行进度。HPIM在两个互补层面运作:进度间全局调制机制在不同执行阶段间自适应未来利用,而进度内相关性机制则在每个进度状态下区分各个未来潜在表征。大量实验表明,与强VLA和WAM基线相比,ProWAM取得了持续的性能提升。
英文摘要
World Action Models (WAMs) extend Vision-Language-Action (VLA) models by incorporating future visual dynamics into action generation. However, existing WAMs often utilize imagined futures with limited adaptation to evolving execution progress, potentially introducing distracting or unreliable predictive cues. This limitation arises from two empirically identified forms of non-uniformity in future utility: (i) at the inter-progress level, the utility of imagined futures varies across execution stages as control demands change; and (ii) at the intra-progress level, individual future latents exhibit heterogeneous relevance within the same progress state. To address these limitations, we propose ProWAM, a Progress-Conditioned World Action Model that introduces execution progress as an explicit intermediate representation for adaptive imagination utilization. ProWAM comprises two tightly coupled components: (1) To obtain a reliable representation of execution progress, we propose the Self-Supervised Dual-Temporal Progress Encoder (SS-DTPE). SS-DTPE couples short-term action-observation interaction modeling with long-term recurrent progress aggregation to capture recent execution feedback and accumulated task history. (2) Conditioned on the progress representation from SS-DTPE, we propose the Hierarchical Progress-Conditioned Imagination Modulation (HPIM) to adapt imagination utilization to execution progress. HPIM operates at two complementary levels: an inter-progress global modulation mechanism adapts future utilization across execution stages, while an intra-progress relevance mechanism differentiates individual future latents within each progress state. Extensive experiments demonstrate consistent gains over strong VLA and WAM baselines.
发表机构
- Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳))
- Great Bay University(大湾区大学)
- University of Electronic Science and Technology of China(电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。