AI 中文总结
本文证明一步式下一潜在预测不能构成世界模型,其误差随展开步数增长,并提出短窗口可改善预测,但各向同性惩罚对转移权重无梯度影响。
AI 中文摘要
下一潜在预测拟合的是从当前嵌入到下一嵌入的映射。LeNEPA 将该目标应用于时间序列,用 LeJEPA 的各向同性惩罚替代下一嵌入预测中的停止梯度。世界模型是可进行展开(rollout)的转移核。一步式回归识别出条件均值,而均值仅在特殊情况下才是核。对于线性高斯马尔可夫潜在变量,一步式问题确定了均值转移和创新协方差,在时间范围 $K$ 上的开环平方误差等于前推创新协方差之和的迹。在一步式拟合精确后,该误差随 $K$ 增长。若条件均值是非线性的,则其复合并非多步条件均值。若观测是马尔可夫状态的非单射函数,则无记忆的一步式映射无法确定未来观测,而短窗口可以。各向同性惩罚是嵌入边际的函数,因此其在转移权重上的偏导数为零。在系数为 $0.9$ 的标量自回归上,一步式均方误差为 $0.998$,16 步开环误差为 $5.10$。在隐藏旋转上,八步窗口达到 16 步误差 $0.056$,而当前标量单独达到 $0.778$。将各向同性权重从 $0.1$ 提高到 $10$,在三个随机种子上八步潜在误差保持在 $[0.78,0.85]$ 内。
英文摘要
Next-latent prediction fits a map from the current embedding to the next one. LeNEPA carries this objective to time series, replacing the stop-gradient of next-embedding prediction with the isotropy penalty of LeJEPA. A world model is a transition kernel that can be rolled out. The one-step regression identifies a conditional mean, and a mean is a kernel only in special cases. For a linear-Gaussian Markov latent, the mean transition and the innovation covariance are fixed by the one-step problem, and the open-loop squared error at horizon $K$ equals the trace of the sum of the pushed-forward innovation covariances. That error grows with $K$ after the one-step fit is exact. If the conditional mean is nonlinear, composing it is not the multi-step conditional mean. If the observation is a non-injective function of a Markov state, a memoryless one-step map does not determine future observations, while a short window can. An isotropy penalty is a function of the embedding marginal, so its partial derivative in the transition weights is zero. On a scalar autoregression with coefficient $0.9$, the one-step mean squared error is $0.998$ and the $16$-step open-loop error is $5.10$. On a hidden rotation, an eight-step window reaches $16$-step error $0.056$, while the current scalar alone reaches $0.778$. Raising the isotropy weight from $0.1$ to $10$ leaves eight-step latent error inside $[0.78,0.85]$ on three seeds.
Comments24 pages, 3 figures