VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
VideoLLM 知何时言:通过视频-文本双人对话交互格式增强时间敏感视频理解
机构 * Wangxuan Institute of Computer Technology, Peking University(王炫计算机技术研究所,北京大学) ; Huawei Noah’s Ark Lab(华为诺亚实验室) ; Beijing Institute for General Artificial Intelligence(北京通用人工智能研究院) ; State Key Laboratory of General Artificial Intelligence(通用人工智能国家重点实验室)
专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV
AI总结 本文提出视频-文本双人对话交互格式,通过MMDuetIT数据集和MAGQA任务提升VideoLLM在时间敏感任务中的表现,实现高效实时响应。
Comments 9 pages