arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OnlineWM:面向有效世界建模的因果感知主动在线学习

OnlineWM: Causality-Aware Active Online Learning for Effective World Modeling

Yikun Miao, Fangqi Zhu, Quanxin Shou, Xiaoyi Pang, Zhengyang Yan, Junhao Li, Haodong Wang, Zicong Hong, Song Guo

arXiv 2609.23753首次发表:更新:

发表机构

The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

OnlineWM提出主动在线学习与因果感知优化框架,通过模拟器交互和反事实学习提升世界模型的动作可控性与泛化能力。

AI 中文摘要

生成式世界模型旨在根据动作预测未来状态,其中动作可控性对于可靠的动力学建模至关重要。尽管最近的工作利用模拟器生成的数据来增强这一能力,但现有的训练流程存在两个根本性局限。首先,静态离线数据收集导致训练集与模型不断演变的错误模式之间存在分布错位,无法解决动力学预测仍不可靠的关键长尾场景。其次,最小化观测差异的标准目标往往鼓励模型利用虚假相关性,而非捕捉潜在的动作-效果因果关系。为解决这些局限,我们提出了OnlineWM,一种通过主动模拟器交互和因果感知优化持续改进世界建模的在线训练框架。OnlineWM引入两项关键创新:(1)主动在线学习:OnlineWM不使用固定数据集,而是自适应地向模拟器查询针对模型当前预测弱点的新交互序列,确保高价值数据获取。(2)因果感知微调:我们提出一种反事实学习策略,对比相同状态下不同动作的结果,迫使模型将状态转变归因于特定动作而非环境自身的演变,从而将其预测建立在可靠的因果机制之上。通过将主动数据获取与因果优化相结合,OnlineWM建立了一个闭环改进过程,确保模型既对多样场景具有鲁棒性,又在因果归因上保持精确。大量实验表明,OnlineWM显著增强了动作可控性,并能有效泛化到未见过的领域。

英文摘要

Generative world models aim to predict future states conditioned on actions, where action controllability is fundamental for reliable dynamics modeling. While recent efforts leverage simulator-generated data to enhance this capability, existing training pipelines face two fundamental limitations. First, static offline data collection leads to a distribution misalignment between training sets and the model's evolving error patterns, failing to resolve critical long-tail scenarios where dynamics predictions remain unreliable. Second, the standard objective of minimizing observational discrepancy often encourages the model to exploit spurious correlations instead of capturing the underlying action-effect causality. To address these limitations, we propose OnlineWM, an online training framework that continuously improves world modeling through active simulator interaction and causality-aware optimization. OnlineWM introduces two key innovations: (1) Active Online Learning: Instead of using fixed datasets, OnlineWM adaptively queries the simulator for new interaction sequences that target the model's current predictive weaknesses, ensuring high-utility data acquisition. (2) Causality-Aware Fine-Tuning: We propose a counterfactual learning strategy that contrasts the outcomes of different actions from identical states, forcing the model to attribute state transitions to specific actions rather than ambient environmental evolution, thereby grounding its predictions in reliable causal mechanisms. By integrating active data acquisition with causal optimization, OnlineWM establishes a closed-loop refinement process that ensures the model is both robust to diverse scenarios and precise in its causal attribution. Extensive experiments demonstrate that OnlineWM significantly enhances action controllability and generalizes effectively to unseen domains.

Comments23 pages, 9 figures. Project page: https://onlinewm.github.io/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑