arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SportD:视觉语言模型能进行物理策略制定吗?

SportD: How do VLMs physically strategize?

Jasin Cekinmez, Addison J. Wu, Haotian Xia, Kyumin Andrew Shim, Anay Putty, Jinglin Xiao, Zhuohan Liu, Leo Liu, Weining Shen

arXiv 2607.14616首次发表:更新:

发表机构

Princeton University; Rice University; UC Irvine; POSTECH; NYU Shanghai; UC Santa Barbara; Zhejiang University(普林斯顿大学; 莱斯大学; 加州大学欧文分校; 浦项科技大学; 纽约大学上海分校; 加州大学圣塔芭芭拉分校; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉语言模型在足球场景中能否进行物理策略制定,引入SportD基准,通过控球价值模型评估模型决策,发现模型表现低于职业球员,有偏好低方差低回报动作等问题,为衡量模型物理策略推理提供测试平台。

AI 中文摘要

视觉语言模型在解释视觉场景方面能力日益增强,但能否利用信息做出策略有效决策尚不明晰。我们在足球场景中研究此问题,模型观察球上决策前几秒,选择射门或传球给特定队友。不同于传统视觉理解任务,足球决策可通过估计每个可用动作的值进行定量评估。我们引入SportD基准,包含2022年世界杯478个球上决策。通过与控球价值模型对比评估模型选择,衡量最优动作准确率和次优决策损失的值。研究发现,三个前沿视觉语言模型最佳表现仅31.4%选最高价值动作,远低于职业球员的38.9%,且都有更大遗憾。进一步分析显示模型偏好低方差、低回报动作,还会部分模仿熟悉模式而非评估反事实替代方案。SportD为衡量视觉语言模型的物理策略推理提供了基于价值的测试平台。

英文摘要

Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whether VLMs can make sound strategic decisions, using soccer as an objective testbed with quantifiably-valued actions. We introduce SportD, a dataset and evaluation consisting of 1421 decision scenarios across professional men's and women's soccer games, where a VLM must decide what action to take next. Models on average select the optimal action around 27% of the time, less often than the professional players, and capture markedly less of the value at stake. Furthermore, they exhibit a clear preference for safer actions, favoring lower-variance, lower-value choices that also make less physical progress toward goal. Frontier VLMs are better at estimating whether an action will succeed, placing the highest-success-probability action among their top choices in 72-85% of cases. Yet VLMs systematically conflate likelihood with value, assigning higher value to actions that are more likely to succeed ($ρ=+0.30$ to $+0.52$), despite no such relationship in the ground truth ($ρ=-0.08$). Modifying the deliberation instructions to encourage risk-taking brings the frontier models closer to the players' skill levels. SportD opens a new direction for rigorously evaluating physical strategic decision-making in VLMs, showing that careful decomposition of their choices can reveal the mechanisms underlying systematic biases such as risk aversion.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑