QCaption: Video Captioning and Q&A through Fusion of Large Multimodal Models
QCaption: 通过融合大多模态模型实现视频描述与问答
机构 * Department of Computer Science and Engineering, Nanyang Technological University, Singapore(南洋理工大学计算机科学与工程系)
专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);分类 cs.CV、cs.AI
AI总结 QCaption通过融合关键帧提取、多模态模型和语言模型,提升视频描述和问答任务的性能,实现44.2%和48.9%的改进。
Journal ref Proceedings of the 27th International Conference on Information Fusion (FUSION), 2024, pp. 1-8