用于潜在世界模型长时序预测的回推解码重构
Rollout-Decoded Reconstruction for Long-Horizon Prediction in Latent World Models
浏览论文内容
中文总结 AI 辅助
本文提出回推解码重构(RDR)损失项,提升了潜在世界模型在混沌Kuramoto-Sivashinsky方程上的长时序预测性能,参数规模不变且仅增加训练开销。
中文摘要 AI 辅助
潜在世界模型会在锚定观测结果的隐变量上训练其解码器,随后将该模型部署在模型自身的自由回推过程中,该过程会延伸至最后一次观测之后的数百步。回推解码重构(RDR)通过一个单一损失项弥合了这一差距,该损失项在训练期间会完全按照评估时的方式自由运行模型,解码每一个回推隐变量,并惩罚与真实值之间的重构误差。该损失项不增加任何参数,仅产生训练时的计算开销,且当权重为零时会简化为标准目标函数,因此本文中的所有对比均为单一标志的A/B测试。在混沌Kuramoto-Sivashinsky方程上,RDR在参数数量同为193568的情况下,将有效预测时间(首次出现归一化误差0.5的时间)从3.87±0.23时间单位提升至6.97±0.42时间单位,提升幅度达1.80倍,该结果在从未用于选择的种子上得到了验证,且在10个预注册配置中,有10个配置的提升比例在1.71至2.50倍之间。这些结果来自单一系统;一项显示优势随隐变量宽度增大而增长的搜索具有描述性,且在两个经典任务上的控制实验为初步结果。
英文摘要
A latent world model trains its decoder on latents anchored to observations, then deploys it on the model's own free-running rollout, hundreds of steps past the last observation. Rollout-Decoded Reconstruction (RDR) closes this gap with a single loss term that free-runs the model during training exactly as evaluation will, decodes every rollout latent, and penalizes reconstruction error against ground truth. The term adds no parameters, costs training-time compute only, and reduces to the standard objective at weight zero, so every comparison in this paper is a one-flag A/B. On the chaotic Kuramoto-Sivashinsky equation, RDR raises valid prediction time (the time to first crossing of normalized error 0.5) from $3.87 \pm 0.23$ to $6.97 \pm 0.42$ time units at an identical 193,568 parameters, a $1.80\times$ improvement confirmed on seeds never used in selection and holding in 10 of 10 preregistered configurations at ratios of 1.71-2.50$\times$. The results come from a single system; a sweep in which the advantage grows with latent width is descriptive, and control experiments on two classic tasks are preliminary.
发表机构
- E3A Healthcare(E3A医疗保健公司)
机构由 AI 辅助整理,请以论文原文为准。