发表机构
Odyssey(奥德赛公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CaliBench用于测试视频世界模型的物理校准,将性能分解为可评分性与校准度,发现多数场景-模型组合显著未校准,发布了mnTV指标用于模型对比。
AI 中文摘要
视频世界模型通过生成式采样近似物理结果的随机分布,但现有基准仅对单个生成结果评分,或在整个数据集上粗略比较分布,未测试特定现象的细粒度偶然不确定性。我们引入CaliBench,它在物理可解释的离散空间(箱索引、骰子面、花色、颜色)中对结果评分,而非像FID那样在学习到的特征空间中进行,因此可直接测量与已知参考分布的距离。我们整理了参考分布以闭式形式已知的结果空间(二项式高尔顿板、伯努利分叉、均匀骰子/纸牌/彩票、偏态欧洲轮盘颜色),从而实现精确校准测试。我们将性能分解为两个正交轴,而单个准确率指标会将其混淆:可评分性(产生可评分结果的生成比例)和校准度(该样本与参考分布的总变差距离)。卡方检验用于评估显著性;由于校准度是其零假设,因此仅能证明未校准,且在每个类别N=32时仅能检测到较大偏差。我们将其应用于9个场景和6个图像转视频模型(WAN-2.7、SeeDance-2.0、HappyHorse-1.0、Veo 3.1、Runway Gen-4.5、Cosmos3-Super),每个模型生成32次结果。模型始终将概率质量集中在少数结果上,而非重现参考分布。大多数场景-模型组合存在显著未校准,极端情况下会崩溃为单个结果,如Veo 3.1在骰子任务上的表现。在轮盘任务中,生成结果常使球的位置模糊,导致多个模型的可评分性较低。性能因场景而异:没有模型在全部9个场景中表现最优。我们发布了该协议和一个指标(平均归一化总变差,mnTV),用于将新模型与我们的结果进行比较。
英文摘要
Video world models approximate the stochastic distribution of physical outcomes through generative sampling, but existing benchmarks score individual generations or compare distributions coarsely over a whole dataset, leaving the fine-grained aleatoric uncertainty of specific phenomena untested. We introduce CaliBench, which scores outcomes in a physically interpretable discrete space - a bin index, a die face, a suit, a colour - rather than a learned feature space such as in FID, so the distance from a known reference distribution is measured directly. We curate outcome spaces whose reference is known in closed form (binomial Galton boards, Bernoulli forks, uniform dice/cards/lottery, a skewed European-roulette colour), enabling an exact calibration test. We decompose performance into two orthogonal axes that a single accuracy metric conflates: scorability, the fraction of generations yielding a scoreable outcome, and calibration, the total variation distance from the reference on that sample. A chi-squared test assesses significance; as calibration is its null hypothesis it can evidence only miscalibration, and at N=32 per cell detects only large deviations. We apply it to nine scenes and six image-to-video models (WAN-2.7, SeeDance-2.0, HappyHorse-1.0, Veo 3.1, Runway Gen-4.5, Cosmos3-Super), 32 generations each. Models consistently concentrate probability mass on a few outcomes rather than reproducing the reference. Most scene-model combinations are significantly miscalibrated, in the extreme collapsing to one outcome, as Veo 3.1 does on dice. On roulette, generations often leave the ball ambiguously placed, giving several models low scorability. Performance varies by scene: no model dominates all nine. We release the protocol and a metric (mean normalised total variation, mnTV) for comparing new models against our results.
CommentsAccepted at Transactions on Machine Learning Research (TMLR). Presented at WOOP @ ECCV 2026. Code is available at https://github.com/odysseyml/calibench
Journal refTransactions on Machine Learning Research 2026