CALIPER:干净场景无法对预训练视觉表征中的物理推理进行排序
CALIPER: Clean Scenes Cannot Rank Physical Inference in Pretrained Visual Representations
浏览论文内容
中文总结 AI 辅助
CALIPER通过先校准后预测的测试证明,干净固定相机场景无法区分预训练视觉表征是否真正推断物理规律,并给出三项检查以验证基准的排序能力。
中文摘要 AI 辅助
被推动物体滑行的距离取决于其质量和摩擦力,而这两者无法从单张图像中获知。预训练视觉编码器越来越多地被用作操作世界模型的感知前端,其物理能力通常通过扰动基准和线性探针进行评估,且几乎总是在干净、固定相机的场景中进行。我们证明,这些评估无法区分能够推断物理规律的编码器与不能推断物理规律的编码器。CALIPER(先校准,后预测)是一种直接测试:一个质量和摩擦力未知的物体以已知速度被击打两次,第三次击打仅展示到接触瞬间,然后在线性冻结特征上使用线性读出器预测物体滑行的距离。通过换用另一个物体的校准片段,可以检验证据是否真正被使用。在2,000个模拟片段和八种表征(从V-JEPA 2到随机初始化的ViT以及原始像素)中,校准增加了+0.50 R²,而换用则消除了这一增益。然而,在干净场景中,每种表征都落在真实模拟器状态所设定的上限的0.02 R²以内,因为固定相机直接在像素坐标中暴露了物体的位移。对每个片段重新采样相机、光照和杂乱背景,使得相同表征的R²分布范围扩大到0.50;当读出器为目标距离选择推动速度时,V-JEPA 2的误差为4毫米,随机ViT的误差为20毫米,与忽略物体相比并无改善。线性探针无法追踪这些差异:帧聚合方式的改变对探针的影响大于预训练本身,而从同一表征中擦除被探测的质量方向,在一个场景中毫无代价,在另一个场景中则损失0.35 R²。基准能否对模型进行排序是一个经验属性,我们提供了三项检查来确立这一点。
英文摘要
How far a pushed object slides depends on its mass and friction, which no single image reveals. Pretrained visual encoders are increasingly used as the perception front end of world models for manipulation, and their physical competence is assessed with perturbation benchmarks and linear probes, almost always in a clean, fixed-camera scene. We show that these assessments cannot distinguish an encoder that infers physics from one that does not. CALIPER (calibrate, then predict) is a direct test: an object of unknown mass and friction is struck twice at known speeds, a third strike is shown only up to the moment of contact, and a linear readout on frozen features must predict how far the object slides. Swapping in another object's calibration clips checks that the evidence is actually used. Across 2,000 simulated episodes and eight representations, from V-JEPA 2 to a randomly initialised ViT and raw pixels, calibration adds +0.50 R^2 and the swap removes it. Yet in the clean scene every representation lands within 0.02 R^2 of the ceiling set by true simulator state, because a fixed camera exposes the object's displacement directly in pixel coordinates. Resampling camera, lighting, and clutter for every clip spreads the same representations across 0.50 R^2; when the readout chooses a push speed for a goal distance, V-JEPA 2 misses by 4 mm and the random ViT by 20 mm, no better than ignoring the object. Linear probes track none of this: a change in frame aggregation moves a probe more than pretraining does, and erasing the probed mass direction from the same representation costs nothing in one scene and 0.35 R^2 in the other. Whether a benchmark can rank models is an empirical property, and we give three checks that establish it.