AI 中文总结
PhysWeep提出一种无标签审计方法,通过黑盒测试视频生成器是否实现请求的物理参数,发现其存在种子条件性错误值收敛的未记录失败,且合理性分数无法察觉。
AI 中文摘要
图像到视频的生成器常被认为能够吸收物理动态,作为隐式世界模型,社区目前通过合理性分数来检验这一说法,即询问片段是否看起来与真实世界运动一致。然而,仅凭合理性是错误的测试,因为一个片段可能在编码了错误的主导物理参数值的同时看起来自然,且现有基准没有直接衡量这一差距。PhysWeep通过一个固定的、无标签的审计来弥补这一差距,将冻结的生成器视为黑盒,从生成的像素中恢复已实现的参数,并报告生成可追踪的频率、已实现值与请求值之间的偏差,以及文献中提出的两种失败机制中哪一种(如果有的话)得到数据支持。一个确定性模拟器的阳性对照确认每个分数都是完全可检验的。应用于三个开放生成器、六个扫描轴,PhysWeep发现了一个具体的、可复现的、先前未记录的失败。在产生可追踪运动的前提下,三个中的两个生成了自信且拟合良好的动态,这些动态收敛到由采样种子而非请求选择的一小组固定错误值之一,并在两个独立的模型家族、两个物理系统和独立追踪器中复现。它既不符合先前的回归,也不符合文献预期的基于案例的钳制,因为回归目标是种子条件性的而非单一全局默认值,且留一法选择规则拒绝了这两者;在可估计响应的任何地方,范围内保真度斜率在统计上与零无法区分。对种子进行平均的基准永远不会看到这一点:每个样本都自信地锁定在一个错误的常数上,这正是合理性分数在结构上无法察觉的失败。我们发布了协议、测试套件和分析代码。
英文摘要
Image-to-video generators are often credited with absorbing physical dynamics as implicit world models, a claim the community currently checks with plausibility scores that ask whether a clip looks consistent with real-world motion. Plausibility is the wrong test on its own, because a clip can look natural while encoding the wrong value of the governing physical parameter, and no existing benchmark measures this gap directly. PhysWeep closes it with a fixed, label-free audit, treating a frozen generator as a black box, recovering the realized parameter from generated pixels, and reporting how often generation is trackable at all, how far the realized value sits from the requested one, and which, if either, of the literature's two proposed failure mechanisms the data support. A deterministic-simulator positive control confirms every score is exactly checkable. Applied to three open generators across six sweep axes, PhysWeep finds a specific, reproducible, previously undocumented failure. Conditional on producing trackable motion, two of the three generate confident, well-fit dynamics that converge to one of a small number of fixed, wrong values selected by the sampling seed rather than by the request, reproducing across two independent model families, two physical systems, and an independent tracker. It matches neither the prior reversion nor the case-based clamping the literature anticipates, because the reversion target is seed-conditional rather than a single global default, and a leave-one-out selection rule rejects both; the in-range faithfulness slope is statistically indistinguishable from zero wherever a response is estimable at all. A benchmark averaging over seeds would never see this: each sample is confidently locked to a wrong constant, exactly the failure a plausibility score is structurally blind to. We release the protocol, suite, and analysis code.