arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30897cs.AI

CAER:面向世界模型训练的因果动作效应重加权

CAER: Causal Action Effect Reweighting for World Model Training

Jianjie Fang, Xvyuan Liu, Ziyou Wang, Rongze Tang, Zhaolu Wang, Zhuohang Li, Xin Zhang, Haisheng Su, Chen Gao, Wei Wu, Xinlei Chen, Yong Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有动作条件世界模型训练中交互动态未充分优化的问题,提出CAER范式,通过在线定位动作因果影响的令牌并重加权,提升了生成视频的物理一致性、可控性与视觉质量。

中文摘要 AI 辅助

世界模型正成为具身智能的核心基础设施,基于动作条件的视频生成可对智能体干预后场景的演化提供可控预测。然而现有模型通常采用时空均匀均方误差(MSE)训练,这使得大量背景令牌主导梯度,而稀疏的交互动态却未得到充分优化;这种均匀拟合奖励外观重构,而非学习动作如何改变世界。我们提出因果动作效应重加权(CAER),这是一种通用训练范式,可将监督重新分配给其预测未来受动作因果影响的令牌。CAER通过对比模型自身在有动作条件和无动作条件下的预测,在线定位这些令牌,随后将所得效应图归一化为权重,该权重保留总系数质量,仅改变权重的分配位置。此在线信号无需外部注释或离线预处理,避免额外数据处理时间,且可随模型和数据集规模自然扩展。在多种异构动作条件世界模型任务上的实验表明,CAER收敛到比均匀MSE训练更好的解决方案,在生成视频的物理一致性、可控性和视觉质量上均有一致提升。

英文摘要

World models are becoming core infrastructure for embodied intelligence, with action-conditioned video generation providing controllable predictions of how scenes evolve after agent interventions. Yet existing models are commonly trained with space-time-uniform mean squared error, allowing abundant background tokens to dominate the gradient while sparse interaction dynamics remain under-optimized; such uniform fitting rewards reconstructing appearance rather than learning how actions change the world. We introduce Causal Action Effect Reweighting (CAER), a general training paradigm that redistributes supervision toward the tokens whose predicted future is causally affected by the action. CAER contrasts the model's own predictions with and without action conditioning to localize these tokens online, then normalizes the resulting effect map into a weight that preserves the total coefficient mass and changes only where it is spent. This online signal requires no external annotations or offline preprocessing, avoids additional data-processing time, and scales naturally with model and dataset size. Experiments across heterogeneous action-conditioned world-model tasks show that CAER converges to better solutions than uniform MSE training, with consistent improvements in the physical consistency, controllability, and visual quality of generated videos.

补充信息

↑