Pinpointing Trigger Moment for Grounded Video QA: Enhancing Spatio-temporal Grounding in Multimodal Large Language Models
机构 * KAIST(韩国科学技术院) ; Ewha Womans University(成均馆大学)
专题命中 视频问答 :video reasoning(abstract);分类 cs.CV
Comments 1st place winner of Grounded Videoqa track at the ICCV2025 Perception Test