arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19354cs.CVcs.AIcs.LG

视觉语言模型能否评判奥运跳水?从推理到零样本动作质量评估的评分

Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment

  • ESPOL Polytechnic University(ESPOL理工大学)
  • University of Granada(格拉纳达大学)
  • Universidad de Las Palmas de Gran Canaria(拉斯帕尔马斯·德·大加那利大学)
  • Michigan Technological University(密歇根理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Henry O. Velesaca, David Freire-Obregon, Luigi Miranda, Abel Reyes-Angulo

AI总结:

本研究提出基于回归的框架,利用开源视觉语言模型对奥运跳水视频进行零样本动作质量评估,通过集成学习将Spearman相关性从低于0.32提升至0.67,证明VLM可作为可解释的半自动体育评估辅助工具。

AI中文摘要:

奥运体育中的自动动作质量评估(AQA)仍是一项具有挑战性的任务,原因在于人体运动的复杂性以及专家评判中固有的主观性。本研究评估了开源视觉语言模型(VLM)在零样本条件下,使用AQA-7基准数据集对奥运跳水视频进行动作质量评估的能力。为此,提出了一种基于回归的框架,利用VLM生成的语义推理和分阶段子评分,结合TF-IDF向量化、降维和集成学习来预测最终比赛得分。实验结果表明,单独的VLM获得的Spearman相关系数适中,低于0.32,而所提出的集成回归框架在报告的评估中显著提升了性能,在四模型配置下达到了0.67的Spearman相关系数。文本推理特征始终优于原始数值子评分,凸显了VLM生成的解释对动作质量分析的丰富性。这些发现表明,VLM作为可解释和半自动体育表现评估的辅助工具具有巨大潜力。代码已在GitHub上公开,可通过此https URL diving judge vlm获取。

英文摘要:

Automated action quality assessment (AQA) in Olympic sports remains a challenging task due to the complexity of human motion and the subjectivity inherent in expert judging. This work evaluates the capability of open-source Vision-Language Models (VLMs) to perform zero-shot action quality assessment on Olympic diving videos using the AQA-7 benchmark dataset. In this regard, a regression-based framework is pro-posed to leverage both the semantic reasoning and phase-level sub-scores generated by the VLMs, combining TF-IDF vectorization, dimensionality reduction, and ensemble learning to predict final competition scores. Experimental results show that standalone VLMs achieve moderate Spearman correlations below 0.32, while the proposed ensemble regression framework substantially improves performance in the reported evaluation, reaching a Spearman correlation of 0.67 with a four-model configuration. Textual reasoning features con-sistently outperformed raw numerical sub-scores, highlighting the richness of VLM-generated explanations for action quality analysis. These findings suggest that VLMs hold strong potential as assistive tools for explainable and semi-automated sports performance evaluation. The code is publicly available on GitHub https://github.com/hvelesaca/olympic diving judge vlm

↑