发表机构
IISER Bhopal; IIT Jodhpur(博帕尔印度科学教育与研究院; 焦特布尔印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CricRAG通过检索相似但更优的技术作为参考,使视觉语言模型提供符合技能水平的个性化板球教练反馈,实现了94%的专业评估一致性。
AI 中文摘要
视觉语言模型(VLMs)为自动化体育教练提供了有前景的能力,但面临一个根本性限制:它们隐式地与专业标准进行比较,这使得它们的反馈对发展中的球员不切实际。我们提出了CricRAG,一个检索增强框架,将视觉语言模型与适合技能水平的基准对齐,用于个性化板球教练。我们的关键见解是,通过检索相似但更优的技术作为参考点,我们可以引导视觉语言模型提供符合发展阶段的反馈,模仿人类教练实践。我们的贡献包括:(1)一个包含288个跨越多个技能水平的板球技术视频的标注数据集,(2)一个使用对比学习的高效动作检索流程,实现了78%的top-3检索准确率,(3)一种降低推理成本的帧采样技术,以及(4)一种检索增强方法,显著提高了反馈与教练原则的一致性,与无检索上下文时的67%相比,实现了高达94%的专业评估一致性。
英文摘要
Vision-Language Models (VLMs) offer promising capabilities for automated sports coaching but face a fundamental limitation: they implicitly compare against professional standards, making their feedback impractical for developing players. We present CricRAG, a retrieval-augmented framework that aligns VLMs with skill-appropriate benchmarks for personalized cricket coaching. Our key insight is that by retrieving similar-but-better techniques as reference points, we can guide VLMs to provide developmentally appropriate feedback that mirrors human coaching practices. We contribute: (1) a labelled dataset of 288 cricket technique videos spanning multiple skill levels, (2) an efficient motion retrieval pipeline using contrastive learning that achieves 78% top-3 retrieval accuracy, (3) a frame sampling technique that reduces inference costs, and (4) a retrieval-augmented approach that significantly improves feedback alignment with coaching principles, achieving up to 94% agreement with professional assessments compared to 67% without retrieval context.
CommentsAAAI 25 - Towards Knowledgeable Foundational Models workshop