arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22788cs.CVstat.AP

人类级准确率,非人类策略:揭示视频物理推理中模型与人类的分歧

Human-Level Accuracy, Non-Human Strategies: Revealing Model-Human Divergence in Video Physical Reasoning

  • Tsinghua University(清华大学)
  • ModelBest Inc.(模本科技公司)
  • Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

Fanhong Li, Shurui Zheng, Zi Yin, Junbo Cui, Lei Ji, Jia Liu

中文总结 AI 辅助

本文提出分布评估框架,揭示视频基础模型在物理推理中虽达人类准确率,但策略与人类系统性分歧,模型依赖统计规律而非向前模拟。

中文摘要 AI 辅助

视频基础模型在物理推理基准上现已达到人类级准确率,然而此类任务要求预测未观测到的物理结果。这些模型是进行类似人类的向前模拟,还是利用可见场景中的统计规律?仅凭准确率无法区分这些策略。我们引入了一个分布评估框架,将模型种子和人类评分者视为总体,从而能够比较共识、不确定性和策略。在Physion基准上,我们评估了三种ViT-L架构(V-JEPA2、VideoMAEv2、DINOv2)。V-JEPA2将准确率差距缩小至约1个百分点(73.2%对比74.2%),然而模型与人类的分歧达到26.4%,远超人类之间的分歧(4.8%),且一致性显著更低(kappa约0.48对比0.91)。这种分歧遵循向前模拟的需求:模型在几何推理(连接,+11.8个百分点)上优于人类,但在引力动力学(滚动,-11.8个百分点)和因果链(多米诺骨牌,-10.5个百分点)上表现不佳。策略指纹识别证实所有三种架构共享非人类策略,而没有一个与人类对齐。归因分析表明,不可观测的结果特征而非可见场景属性预测了这种分歧,这与模型更多依赖场景级统计规律而非显式向前模拟的结论一致,这是一种仅凭准确率无法揭示的系统性分歧。代码可在该https URL获取。

英文摘要

Video foundation models now reach human-level accuracy on physical-reasoning benchmarks, yet such tasks require predicting unobserved physical outcomes. Do these models perform human-like forward simulation, or do they exploit statistical regularities in visible scenes? Accuracy alone cannot distinguish these strategies. We introduce a distributional evaluation framework that treats model seeds and human raters as populations, enabling comparison of consensus, uncertainty, and strategy. On the Physion benchmark, we evaluate three ViT-L architectures (V-JEPA2, VideoMAEv2, DINOv2). V-JEPA2 narrows the accuracy gap to ~1 percentage point (73.2% vs. 74.2%), yet model-human disagreement reaches 26.4%, far exceeding human-human disagreement (4.8%), with substantially lower agreement (kappa ~ 0.48 vs. 0.91). The divergence follows forward-simulation demands: models outperform humans on geometric reasoning (linking, +11.8 pp) but underperform on gravitational dynamics (rolling, -11.8 pp) and causal chains (dominoes, -10.5 pp). Strategy fingerprinting confirms all three architectures share non-human strategies while none aligns with humans. Attribution analysis suggests that unobservable outcome features, rather than visible scene properties, predict this divergence, consistent with models relying more on scene-level statistical regularities than on explicit forward simulation, a systematic divergence that accuracy alone cannot reveal. Code is available at https://github.com/fanhong-li/model-human-divergence.

补充信息

相关深度报道

↑