arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.28406cs.CVcs.LG

用于视频错误检测的后训练视觉语言模型

Post-Training VLMs for Video Mistake Detection

Federico Spurio, Olga Zatsarynna, Lars Doorenbos, Emad Bahrami, Gianpiero Francesca, Juergen Gall

首次发表
浏览论文内容

中文总结 AI 辅助

针对视频错误检测的闭集方法泛化性差的问题,提出首个VLM后训练技术,采用定制奖励函数,在EP-VQA上较最佳基线提升11.6%,实现更好的泛化。

中文摘要 AI 辅助

在遵循指令时,人为错误不可避免,且可能导致严重后果。因此,开发视频错误检测方法的兴趣日益浓厚,当前方法大多聚焦于闭集协议。尽管在受控环境中表现成功,但闭集假设限制了其更广泛的适用性,因为任务的任何变化都需要收集新数据并重新训练模型。相反,我们认为错误检测方法应学习错误的通用概念,而非过拟合于特定步骤的细节。为体现这一点,我们引入了错误检测视频问答(Mistake Detection Video Question Answering,MD-VQA)协议及配套基准。MD-VQA测试方法能否针对已见和未见动作,判断某一步骤是否相对于其描述被正确执行。为应对这一重要挑战,我们提出了首个用于错误检测的视频语言模型(Video-Language-Model,VLM)后训练技术。我们的方法采用定制奖励函数,鼓励模型识别指令与对应视频之间的差异。大量评估表明,该方法优于零样本、监督微调及后训练基线。值得注意的是,我们的方法对未见流程的泛化效果尤为出色,例如在EP-VQA上较表现最佳的基线提升了11.6%,为通用错误检测铺平了道路。我们在该httpsURL发布代码和基准。

英文摘要

Human mistakes are inevitable when following instructions, yet they can lead to severe consequences. As such, there has been an increased interest in developing methods for detecting mistakes in videos, with current methods mostly focusing on closed-set protocols. While successful in controlled settings, the closed-set assumption limits their wider applicability, as any changes to the task require collecting new data and re-training models. Instead, we argue that mistake detection methods should learn the general concept of a mistake, rather than overfitting to step-specific details. To reflect this, we introduce the Mistake Detection Video Question Answering (MD-VQA) protocol and accompanying benchmark. MD-VQA tests whether methods can discern if a step was executed correctly with respect to its description, for both seen and unseen actions. To address this important challenge, we propose the first video-language-model post-training technique for mistake detection. Our method uses a tailored reward function to encourage the model to identify discrepancies between an instruction and the corresponding video. Extensive evaluations demonstrate that this approach outperforms zero-shot, supervised fine-tuning, and post-training baselines. Notably, our method generalizes especially well to unseen procedures, for instance, with an improvement of up to 11.6% over the best-performing baseline on EP-VQA, paving the way toward general mistake detection. We release our code and benchmark at https://github.com/FedeSpu/mstk.

发表机构

  • University of Bonn(波恩大学)
  • Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔机器学习与人工智能研究所)
  • Toyota Motor Europe Belgium(丰田汽车欧洲比利时公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑