arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32657cs.AI

具有可预测长时域边际的世界模型

World Models with Predictable Long-Horizon Marginals

Yuhao Du, Shunian Chen

AI总结:

本文提出一种通过固定高斯参考解码器与条件动作图约束转移的世界模型,实现长时域展开保持数据分布,并在36个任务-种子单元中于10^5步内保留所有评估链。

AI中文摘要:

准确的一步预测并不能确保世界模型的展开(rollouts)保持数据分布。我们通过学习一个固定高斯参考的解码器,并约束行为平均的转移以保持该参考,从而使模型的解码稳态分布显式化。对于受控系统,联合转移使用条件动作图(conditional action chart)来保持行为占用率,而无需在每个固定动作下要求不变性。联合状态-动作旋转和平行高斯噪声给出了一个精确保持的转移,并具有可处理的条件密度。我们从有限初始化库推导出绝对收敛界,并通过条件动作空间散度控制与参考的偏离。在十二个控制任务的216个拟合像素检查点上,具有参考混合的占用模型在所有36个任务-种子单元中,于10^5步内保留了每条评估链,滚动减去参考的能量统计差异为-0.0002±0.0003(训练种子标准误)。四个非保持对照臂均丢失了链,尽管高斯臂在十步时更准确。一个离线的DreamerV3参考也实现了更好的短时域准确性。这些结果区分了世界模型的三个属性:其逼近的分布、逼近的速率以及其学习的条件动力学。

英文摘要:

Accurate one-step predictions do not ensure that a world model's rollouts retain the data distribution. We make the model's decoded stationary law explicit by learning a decoder of a fixed Gaussian reference and constraining the behaviour-averaged transition to preserve that reference. For controlled systems, a joint transition uses a conditional action chart to preserve behaviour occupancy without requiring invariance at each fixed action. Joint state--action rotations and parallel Gaussian noise give an exactly preserving transition with a tractable conditional density. We derive an absolute convergence bound from finite initialization banks and control departure from the reference through conditional action-space divergence. Across $216$ fitted pixel checkpoints on twelve control tasks, the occupancy model with a reference mixture retains every evaluated chain at $10^5$ steps in all $36$ task--seed cells, with a rollout-minus-reference energy-statistic difference of $-0.0002\pm0.0003$ (training-seed standard error). Each of the four nonpreserving comparison arms loses chains, although the Gaussian arm is more accurate at ten steps. An offline DreamerV3 reference also achieves better short-horizon accuracy. These results distinguish three properties of a world model: the distribution it approaches, the rate of approach, and the conditional dynamics it learns.

↑