Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
机构 * University of California, Davis(加州大学戴维斯分校) ; Virginia Tech(弗吉尼亚理工大学) ; The Chinese University of Hong Kong(香港中文大学) ; NVIDIA ; Adobe Research(Adobe研究) ; Fudan University(复旦大学) ; Meta AI
专题命中 视频理解 :video understanding(abstract);分类 cs.CV
Comments Accepted by EMNLP 2025 Findings