arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23369cs.CV

行动的正确未来:在世界行动模型中学习行动相关的预测状态

The Right Future for Action: Learning Action-Relevant Predictive States in World Action Models

Qiwen Gu, Jifan Li, Bingjie Gao, Rui Chen, Jing Tang, Xiangxiang Chu, Junqiao Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出行动相关预测状态(ARPS),通过未来表示监督学习紧凑预测状态,提升世界行动模型在分布偏移下的泛化,在LIBERO和LIBERO-Plus上分别达到99.2%和87.3%的成功率。

中文摘要 AI 辅助

免生成世界行动模型(WAMs)在训练期间保留未来视频预测,但在推理时从内部视频特征中行动,这留下了这些特征应为控制保留什么的问题。我们的表示诊断表明,具有更可预测的未来变化的表示不一定使线性行动解码更容易。观察到的未来变化提供了超越现在的额外行动信息,且线性可读的行动信息在空间上集中。这些发现激发了行动相关预测状态(ARPS),这是视频专家和行动专家之间的紧凑预测接口。ARPS使用一个以地平线为条件的状态预测器,将中间视频特征聚合为一个紧凑状态,为行动专家提供所有视觉上下文。未来表示监督训练该状态的不同部分,以预测不同未来时间的视觉表示及其相对于现在的变化。在推理时,监督分支被移除,行动专家仅使用从当前观测计算出的学习预测状态。受控消融表明,未来监督显著提高了分布偏移下的泛化能力。ARPS在LIBERO上达到99.2%的成功率,并无需适应即可迁移到LIBERO-Plus,达到87.3%,超过Fast-WAM 39.2个百分点。

英文摘要

Generation-free world action models (WAMs) retain future-video prediction during training but act from internal video features at inference, leaving unclear what these features should preserve for control. Our representation diagnostics show that representations with more predictable future changes need not make linear action decoding easier. Observed future changes provide additional action information beyond the present, and linearly readable action information is spatially concentrated. These findings motivate Action-Relevant Predictive States (ARPS), a compact predictive interface between the video and action experts. ARPS uses a horizon-conditioned state predictor to aggregate intermediate video features into a compact state that supplies all visual context to the action expert. Future-representation supervision trains different parts of this state to predict visual representations at different future times, together with their changes relative to the present. At inference, the supervision branch is removed, and the action expert only uses the learned predictive state computed from current observations. Controlled ablations show that future supervision substantially improves generalization under distribution shift. ARPS achieves 99.2% success on LIBERO and transfers to LIBERO-Plus without adaptation, reaching 87.3% and exceeding Fast-WAM by 39.2 percentage points.

发表机构

  • Tongji University(同济大学)
  • DreamX Team, Alibaba Group(阿里巴巴集团DreamX团队)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑