arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11174cs.RO

VIScore:诊断潜在世界模型中与规划相关的质量

VIScore: Diagnosing Planning-Relevant Quality in Latent World Models

发表机构阿尔托斯实验室 · 布朗大学
查看机构详情
  • Altos Labs(阿尔托斯实验室)
  • Brown University(布朗大学)

机构由 AI 辅助整理,请以论文原文为准。

Haiyu Wu, Randall Balestriero, Morgan Levine

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对潜在世界模型规划成功率与潜在空间属性脱节的问题,提出VIScore指标,覆盖编码器、预测器、规划器,在跨任务场景下斯皮尔曼相关性超0.75,可更好解释规划成功率。

中文摘要 AI 辅助

将潜在空间正则化为各向同性高斯分布,可为世界模型规划提供稳定且信息最大化的空间。然而,潜在空间属性与成功规划之间仍存在脱节。我们通过比较SIGReg和VISReg两种正则化损失函数来研究这一问题,二者具有相同的分布目标但属性不同。与SIGReg相比,VISReg在控制中心、尺度和形状正则化的权重方面具有更高的灵活性,且更大的批量大小可实现更精细的分布近似。我们发现,前者虽在自监督学习(SSL)中有益,但对规划无帮助;而后者可提高分布外(OOD)数据集上的规划成功率。这促使我们深入研究与成功率相关的因素。与仅关注编码潜在空间的现有指标不同,我们提出了真实性-影响力-清醒度分数(VIScore),该指标可量化给定编码特征的预测器的可达性和容量,以及基于搜索的规划器的幻觉。与直线度、物理状态探测和赋能相比,我们证明,由于VIScore的测量覆盖了编码器、预测器和规划器,其比其他指标更能解释成功率,表现为强斯皮尔曼相关性。具体而言,在跨任务成功率池中,VIScore在已见和未见模型及数据集上均始终实现超过0.75的斯皮尔曼相关性。此外,VIScore是唯一在所有测试场景中校准误差低于常数拟合的指标,凸显了这三个方面对规划成功的重要性。我们希望该指标能助力未来世界模型的设计与诊断研究。

英文摘要

Regulating the latent space to an isotropic Gaussian distribution provides a stable and information-maximized landscape for world model planning. However, the latent space property and successful planning remain disconnected. We first study this by comparing SIGReg and VISReg, two regularization loss functions with the same distribution target but different properties. Compared with SIGReg, VISReg has more flexibility in controlling the weights of center, scale, and shape regularization, and a larger batch size brings a finer distribution approximation. We find that the former, despite being beneficial in self-supervised learning (SSL), does not help the planning, whereas the latter improves the planning success on out-of-domain (OOD) datasets. This motivates a deep understanding of the factors that correlate with the success rate. Unlike the previous metrics focusing on the encoded latent only, we propose the Veracity-Influence-Sobriety score (VIScore), a metric that quantifies the reachability and capacity of a predictor given the encoded feature, and the hallucination of the searching-based planner. Compared with straightness, physical-state probing, and empowerment, we show that, with the measurement covering encoder, predictor, and planner, VIScore explains the success rate better than the others, as reflected by a strong Spearman correlation. Specifically, VIScore consistently achieves a Spearman correlation over 0.75 on both seen and unseen models and datasets on the cross-task success rate pool. Moreover, VIScore is the only metric that has a calibration error below the constant fit across all testing scenarios, showcasing the importance of these three aspects in planning success. We hope this metric can help future studies on world model design and diagnosis.

补充信息

↑