基于文本描述和合成图像的零样本视频高光检测
Zero-shot video highlight detection based on text descriptions and synthetic images
浏览论文内容
中文总结 AI 辅助
提出一种结合CLIP、大语言模型和扩散模型的零样本框架,利用视频元数据生成文本描述和合成图像,实现无需标注的帧级高光检测,在TVSum和SumMe上表现优异。
中文摘要 AI 辅助
检测视频高光,即视频中最具信息量或最吸引人的时刻,对于视频摘要和内容推荐等应用非常重要。我们提出了一种结合CLIP、大型语言模型(LLMs)和扩散模型的零样本框架。给定轻量级视频元数据(如标题或类别),LLM生成可能高光事件的文本描述。这些描述进一步通过扩散模型转换为合成视觉原型。文本和视觉表示通过CLIP与视频帧进行匹配,从而在没有高光标注或特定数据集训练的情况下实现帧级高光检测。在TVSum和SumMe上的实验展示了强大的零样本性能,在TVSum上结果尤为突出。所提出的方法为基于元数据条件的零样本视频高光检测提供了一个有效的框架。
英文摘要
Detecting video highlights, the most informative or engaging moments in a video, is important for applications such as video summarization and content recommendation. We propose a zero-shot framework that combines CLIP, large language models (LLMs), and diffusion models. Given lightweight video metadata, such as a title or category, an LLM generates textual descriptions of likely highlight events. These descriptions are further converted into synthetic visual prototypes using a diffusion model. Textual and visual representations are matched to video frames using CLIP, enabling frame-level highlight detection without highlight annotations or dataset-specific training. Experiments on TVSum and SumMe demonstrate strong zero-shot performance, with particularly favorable results on TVSum. The proposed approach provides an effective framework for metadata-conditioned zero-shot video highlight detection.
发表机构
- Samsung AI Center(三星人工智能中心)
- Institute of Fundamental Technological Research, Polish Academy of Sciences(波兰科学院基础技术研究所)
机构由 AI 辅助整理,请以论文原文为准。