arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08162cs.AI

WorldAgen:统一状态-动作预测与测试时世界模型训练

WorldAgen: Unified State-Action Prediction with Test-Time World Model Training

Chi Wan, Kangrui Wang, Yuan Si, Pingyue Zhang, Manling Li

首次发表
浏览论文内容

中文总结 AI 辅助

WorldAgen提出统一框架,联合世界建模与动作预测,并通过测试时训练适应新环境,在CALVIN和LIBERO基准上超越现有最先进方法。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型如何适应世界动态发生变化的新环境?尽管近期研究将世界建模与动作预测相结合以提升VLA性能,但现有方法大多依赖在静态数据集上的预训练,缺乏部署时主动适应的机制。因此,这些模型在部署到具有新颖物体配置或动态的未见场景时,往往无法泛化。我们提出WorldAgen,一个统一框架,联合学习世界建模与动作预测,并支持测试时训练(TTT)以适应新环境。WorldAgen采用共享Transformer主干,配备两个头:(1)世界模型头,根据过去的状态-动作轨迹预测未来状态;(2)智能体模型头,根据任务指令预测动作。我们设计了一种混合单向注意力掩码来分离这两个模型。在测试时,WorldAgen采样探索性动作,收集真实状态转移,并执行轻量级TTT更新以精炼其世界模型。这种适应提升了模型对环境理解,并带来更准确的动作预测。在CALVIN和LIBERO基准上的实验表明,我们的基线模型达到了与当前最先进方法相当甚至更优的性能。此外,在少量样本上进行TTT后,我们的方法超越了现有最先进模型,突显了推理时适应世界模型的有效性。

英文摘要

How can vision-language-action (VLA) models adapt to new environments where world dynamics shift? While recent research has combined world modeling and action prediction to improve VLA performance, existing methods largely rely on pretraining on static datasets, without mechanisms for active adaptation at deployment time. As a result, these models often fail to generalize when deployed in unseen scenarios with novel object configurations or dynamics. We present WorldAgen, a unified framework that jointly learns world modeling and action prediction while enabling Test-Time Training (TTT) to adapt to new environments. WorldAgen employs a shared Transformer backbone with two heads: (1) a world model head that predicts future states from past state-action trajectories, and (2) an agent model head that predicts actions conditioned on task instructions. We design a Mixed Unidirectional Attention Mask to separate these two models. During test time, WorldAgen samples exploratory actions, collects ground-truth state transitions, and performs lightweight TTT updates to refine its world model. This adaptation improves the model's understanding of the environment and leads to more accurate action predictions. Experiments on the CALVIN and LIBERO benchmarks demonstrate that our baseline model achieves comparable, and in some cases superior, performance to current state-of-the-art approaches. Moreover, with TTT on a small number of samples, our method surpasses existing state-of-the-art models, highlighting the effectiveness of adapting world models at inference time.

发表机构

  • Northwestern University(西北大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑