HAS:用于多模态大语言模型视频摘要的高亮引导注意力转向
HAS: Highlight-guided Attention Steering for Multimodal LLM Video Summarization
浏览论文内容
中文总结 AI 辅助
针对视频摘要,提出HAS方法,通过全局找连续帧级高亮分布并作为注意力转向向量,让多模态大语言模型在推理时更关注高亮帧,避免丢失信息,在多种基准测试中性能出色。
中文摘要 AI 辅助
随着人工智能在视频生成方面的发展,视频理解变得越来越重要。多模态大语言模型已展现出视频理解能力,视频摘要作为视频理解的特定领域,对高效导航和检索很重要。当前视频摘要方法侧重于选定关键帧和相关片段字幕,忽视了全局帧重要性。本文提出HAS方法,考虑视频连续但有高亮引导的实际情况,由全局找连续帧级高亮分布和将其作为注意力转向向量两部分组成,在多个基准测试中展现出令人信服的性能。
英文摘要
Video understanding has become more and more important with the growth of Artificial Intelligence (AI) for video generation. Recently, Multimodal Large Language Model(M-LLM) has shown its capability in video understanding. Video summarization, a specific domain of video understanding, has proven its importance for efficient navigation and retrieval. Both video understanding and video summarization require a good selection of key frames in a video. Current video summarization methods heavily focus on the selected key frames and correlated segment captions. However, existing approaches overlook the perspective of treating the importance of the frames globally. We argue that using discrete selected frames for summarization will not only reduce the understanding coherence, but also lost important information in the video, as well as wasting the original capacity of the MLLMs. In this paper, we propose HAS, a Highlight-guided Attention Steering method for video summarization. We consider a challenging but practical setting where the video given to MLLMs for summarize should be continuous but with highlight guidance. HAS mainly consists of two parts: The first part is to find a continuous frame-level highlight distribution for the video globally. The second part is to apply the highlight distribution as an attention steering vector for the MLLM, targeting a better understanding of the video, and thus during the model inference time, putting more attention on the highlighted frames, while avoiding lost entire information on less highlighted frames through putting less attention instead of forgetting them. We evaluated HAS on a variety of benchmarks, and it has shown convincing performance in video summarization.
发表机构
- Tufts University(塔夫茨大学)
机构由 AI 辅助整理,请以论文原文为准。