arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

算子对F补充值等价性:潜在世界模型的规划时诊断

Operator-on-F complements value-equivalence: a planning-time diagnostic for latent world models

Donna Vakalis

arXiv 2607.04464首次发表:更新:

AI 中文总结

研究针对基于模型的强化学习中世界模型评估问题,引入算子对F诊断,通过模型自身预测器比较模型与环境的k步潜在前推,揭示其与规划回报的强关联及在跨架构比较中的作用,补充而非替代值等价性诊断。

AI 中文摘要

基于模型的强化学习中的世界模型评估通常询问学习到的模型是否能很好地预测奖励和价值,这可能使模型潜在展开中与规划相关的错误未被测量。我们引入一种补充诊断,算子对F,它使用模型自己的预测器在可观察子集F上比较模型的k步潜在前推与环境的。在对猎豹奔跑的TD-MPC2大小扫描中,奖励预测误差在每个模型大小下都保持在[0.028, 0.091]内,只有约3倍的变化,因此未归一化的奖励拟合检查区分它们的分辨率很窄;(未归一化的)贝尔曼残差和奖励误差本身与回报的关系较弱(斯皮尔曼相关系数为-0.10和-0.30)。在相同大小范围内,算子误差跨度为0.28至2.62。在317M时,算子误差为2.62,比0.28 - 0.36的聚类高出一个数量级,规划回报降至0.9,而奖励预测误差(0.091)是五个中最高的,但仍与扫描的其余部分保持在相同的小[0.028, 0.091]范围内。算子误差与回报损失之间的秩相关系数为-0.90(在n = 5个大小下的锚定自举95%置信区间[-0.90, -0.70];去除任何单个大小的留一法使其保持在-0.80或更强)。该算子还在TD-MPC2和纯SSL潜在世界模型的跨架构比较中返回信息丰富、区分架构的估计。算子诊断补充值等价性而非取代它。

英文摘要

World-model evaluation for model-based reinforcement learning typically asks whether the learned model predicts reward and value well, which can leave planning-relevant errors in the model's latent rollouts unmeasured. We introduce a complementary diagnostic, operator-on-F, that compares a model's k-step latent pushforward to the environment's on an observable subset F, using the model's own predictor. On a TD-MPC2 size sweep over cheetah-run, reward-prediction error stays within [0.028, 0.091] for every model size - only about 3x variation - so an unnormalized reward-fit check has narrow resolution to distinguish them; the (unnormalized) Bellman residual and reward error themselves have weak relationships with return (Spearman -0.10 and -0.30). Operator error spans 0.28 to 2.62 over the same sizes. At 317M the operator error is 2.62 - an order of magnitude above the 0.28-0.36 cluster - and the planning return collapses to 0.9, while reward-prediction error (0.091) is the highest of the five but stays within the same small [0.028, 0.091] range as the rest of the sweep. The rank correlation between operator error and return loss is -0.90 (anchor-bootstrap 95% CI [-0.90, -0.70] at n=5 sizes; leave-one-out removal of any single size leaves it at -0.80 or stronger). The operator also returns informative, architecture-discriminating estimates in a cross-architecture comparison between TD-MPC2 and a pure-SSL latent world model. The operator diagnostic complements value-equivalence rather than replacing it.

CommentsAccepted at RLC 2026 WM Workshop. V2 places the diagnostic in Koopman representation theory; references expanded. Results unchanged

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑