未来为何分支?随机物理世界模型的可识别闭合测试
Why Does the Future Branch? Identifiable Closure Tests for Stochastic Physical World Models
浏览论文内容
中文总结 AI 辅助
针对随机物理世界模型无法从普通转移数据识别未来分支原因的问题,提出ClosurePairs干预评估协议,可精准分解状态混淆与过程噪声,大幅提升归因准确率。
中文摘要 AI 辅助
随机世界模型通常通过其预测未来的准确性和校准度进行评估。这些标准留下了与决策相关的歧义:相同的条件未来分布可能源于观测混淆了不同的物理状态,也可能源于在声明完整状态固定后动力学仍保持随机性。我们证明,即使使用最优概率预测器,也无法从普通转移数据中识别这种归因。我们引入ClosurePairs,一种将兼容微态与重复外生扰动交叉的干预评估协议。双向方差分解可识别状态混淆、过程噪声及其非线性相互作用;当无法重复使用扰动时,可采用独立重复变体。在似然等价高斯系统中,配对监督在相同测试负对数似然(NLL)下将混淆分数误差降低15.96倍;在18种非线性朗之万条件下,其在不改变NLL的情况下将归因平均绝对误差(MAE)从0.372降至0.051,感知遗憾从0.0138降至0.0003。在像素条件循环模型上,冻结的共享状态探针在分布内将混淆分数MAE从深度集成的0.584降至0.130,在分布外十次随机种子实验中从0.630降至0.170。最后,在匹配总方差的REFINE/BRANCH测试中,总方差路由器达到66.48%±1.06%的准确率,而ClosurePairs达到99.99%±0.02%的准确率,并在五次随机种子实验中将选定NLL从-2.087提升至-2.717。因此,ClosurePairs可测量未来分支的原因,这是 proper forecast scores 无法识别的信息。
英文摘要
A calibrated stochastic world model can reveal how uncertain a future is without revealing why it branches. The same conditional future law can arise because an observation aliases physical states or because dynamics remain random after the declared full state is fixed. We prove that ordinary transitions cannot identify these two sources, even for a perfect probabilistic predictor. ClosurePairs makes them identifiable by crossing compatible microstates with repeated exogenous disturbances and estimating state, noise, and state-noise interaction variance. The central consequence is operational: under finite hierarchical sampling, forecast difficulty governs the useful compute scale, while the alias/process composition provides complementary information about its direction-resolving the current state or sampling future randomness. ClosurePairs recovers source attribution at unchanged likelihood, reduces equal-budget decomposition error in a nonlinear interaction benchmark, and supports observation-only routing. On exact-marginal MetaWorld twins, an output-only allocator is at chance while a Closure-supervised probe on frozen JEPA-WM features routes 89.8-100%. In an independent ManiSkill PushCube confirmation, a stochastic RSSM's outputs and latents remain at chance, whereas an RGB-only Closure probe routes 100% under both ID and geometry/camera OOD over five seeds, matching direct allocation rather than exceeding it. Across five unseen allocation menus, the same Closure probe routes 92.5%/90.4% ID/OOD with no new oracle labels, versus 37.9%/32.9% for a frozen direct allocator. ClosurePairs is therefore an identifiable, reusable mechanism target that cannot be recovered from forecast quality alone.
发表机构
- Shandong University(山东大学)
机构由 AI 辅助整理,请以论文原文为准。