Post-Training VLMs for Video Mistake Detection
用于视频错误检测的后训练视觉语言模型
机构 * University of Bonn(波恩大学) ; Lamarr Institute for Machine Learning and Artificial Intelligence(拉马尔机器学习与人工智能研究所) ; Toyota Motor Europe Belgium(丰田汽车欧洲比利时公司)
专题命中 视觉问答 :VLM(summary_cn,abstract_cn);分类 cs.CV、cs.LG
AI总结 针对视频错误检测的闭集方法泛化性差的问题,提出首个VLM后训练技术,采用定制奖励函数,在EP-VQA上较最佳基线提升11.6%,实现更好的泛化。
Comments Accepted at BMVC 2026